MobileForge in Practice: Putting GUI Agents on an Unlabeled Data Flywheel
Kuaishou, together with Zhejiang University, open-sources MobileForge: MobileGym + HiFPO let mobile GUI agents self-explore, self-feed-back, and self-improve inside real apps, with full code and models released.


MobileForge in Practice: Putting GUI Agents on an Unlabeled Data Flywheel
Kuaishou, together with Zhejiang University, open-sources MobileForge: MobileGym + HiFPO let mobile GUI agents self-explore, self-feed-back, and self-improve inside real apps, with full code and models released.
The biggest bottleneck for mobile GUI agents on real apps isn't "can't tap" — it's "can't adapt." Apps number in the millions and update constantly, and adapting each one means hand-writing tasks, recording expert trajectories, and labeling reward signals, so costs spiral out of control. MobileForge, open-sourced by Kuaishou together with Zhejiang University and Tsinghua, turns this into an unlabeled closed loop: the agent self-explores, self-feeds-back, and self-improves inside real apps. This tutorial walks you through the full pipeline.
What MobileForge Is
MobileForge consists of two coupled components:
- MobileGym: the interaction and evaluation foundation. It explores reachable states in the target app, mines executable tasks from real interaction trajectories, and runs hierarchical evaluation over the complete execution process
- HiFPO (Hierarchical Feedback-Guided Policy Optimization): hierarchical feedback-guided policy optimization. It schedules multiple attempts, reuses hints from failures, filters for valuable tasks and steps, and updates the model with hint-contextualized step-level GRPO
The whole pipeline uses no hand-written tasks, no expert demonstrations, and no human reward labels:
Target app exploration → task curriculum generation → multiple rollouts → hierarchical evaluation → task/trajectory/step filtering → GRPO training with corrective hints
Before You Start
Resource List
- Paper: https://arxiv.org/abs/2606.19930
- Project page: https://mobile-forge.github.io/
- GitHub: https://github.com/kwai/MobileForge
- Full-pipeline datasets: https://huggingface.co/collections/lgy0404/mobileforge-datasets
- Full-pipeline models: https://huggingface.co/collections/lgy0404/mobileforge-models
Measured Results (Reference)
| Model | AndroidWorld Pass@3 | Notes | | | | | | Qwen3-VL-8B (baseline) | 55.2% | General-purpose VLM | | ForgeQwen3-8B (after adaptation) | 67.2% | Approaches GUI-specialized foundations | | GUI-Owl-1.5-8B (baseline) | 69.0% | GUI-specialized | | ForgeOwl-8B (after adaptation) | 77.6% | Strongest in-domain |
Cross-domain (MobileWorld GUI-only, with no MobileWorld data seen during training): ForgeOwl-8B reaches 41.0%, above the baseline's 37.6%.
What You Need
- A GPU machine capable of training an 8B VLM (multi-GPU is better; multi-node training scripts are provided)
- Familiarity with PyTorch / HuggingFace transformers
- An Android emulator or real-device environment (for AndroidWorld rollouts)
- Python 3.10+
Step 1: Get MobileGym Exploration Running
MobileGym answers "what should the agent learn when there are no hand-written tasks?" It has three stages.
1.1 Target App Exploration
MobileForge enters the target app directly and, combining structural information such as the activities declared in the APK with the current screenshot, generates feature-oriented exploration goals. Exploration uses a depth-first-like traversal; when it needs to branch from a parent state to a new goal, it restores the parent state and continues from there.
Every explored state transition is recorded: before/after screenshots, the executed action, the target element, execution metadata, and a natural-language summary. These records form the evidence pool.
Note: exploration trajectories are not expert demonstrations. Their sole purpose is to discover genuinely reachable screens, operable controls, and features that actually exist, keeping the model from hallucinating out of thin air.
1.2 MobileGym-Curriculum: Turning Evidence into Tasks
For each exploration trajectory, the system judges whether the behavior was coherent and whether the original goal was achieved, then generates multiple task variants around the same app feature.
Each task is a five-tuple: (task instruction, estimated step budget, core feature, change type, preconditions). The point isn't a complex schema — it's that every task must be anchored to genuinely observed app behavior.
1.3 MobileGym-Critic: Hierarchical Evaluation
The Critic is not a trained reward model; an agentic hierarchical evaluator produces three types of feedback on the complete rollout:
- Trajectory-level outcome label: whether the task was ultimately completed
- Step-level process label: whether each step was reasonable, and why
- Corrective hint: summarizes why it failed, behaviors to avoid, suggested alternative strategies, and key task insights
This step is what sets MobileForge apart from conventional RL: failed trajectories contain correct local steps, and successful trajectories contain redundant actions — the Critic pulls this information apart.
Step 2: Run the HiFPO Training Loop
HiFPO turns MobileGym's feedback into policy updates in four steps.
2.1 Multiple Attempts with Hints
Each task is attempted K times in a row by the current policy:
- The first attempt gets no extra hint
- On failure or unreasonable behavior, the Critic generates a corrective hint
- On the second attempt, the hint is appended to the task instruction
The agent isn't just sampling more; it accumulates experience on the same task — the previous failure becomes the next attempt's context. Measured (Qwen3-VL-8B, 200 tasks): overall success rate 52.0% without hints → 77.0% with hints, and Pass@3 from 49.0% → 72.5%.
2.2 Task Filtering
Compute each task's empirical success rate SR(x) over its multiple attempts:
- All-success tasks: already mastered by the current policy; removed (little training value)
- All-failure / partial-success tasks: kept
This is counterintuitive: MobileForge doesn't discard failed tasks, because failed trajectories may contain correct navigation/search/recognition steps — as long as step-level feedback can pick them out, failures can be converted into learning material too.
2.3 Trajectory and Step Selection
For the kept tasks:
- With successful trajectories: pick the successful trajectory with the highest step quality
- All failed: pick the failed trajectory with the highest proportion of locally reasonable steps
- The training set keeps only the local steps the Critic judged reasonable
Long-horizon trajectories get split into dense step-level samples, while avoiding reinforcing the wrong actions inside failed trajectories.
2.4 Hint-Contextualized Step-Level GRPO
Each step-level sample contains the task, screenshot, interaction history, and the corrective hint at that moment. The model samples multiple candidate actions under the same hint-conditioned state, and group-relative comparison (GRPO) is done with a rule-based GUI action reward.
This is the training-objective ablation finding from MobileForge: no-hint SFT performs weakly (even below baseline), hint SFT improves things, but hint-contextualized GRPO is best in both the 200- and 900-task settings. At 900 tasks it reaches 50.9% AndroidWorld Pass@1.
Step 3: Evaluate
In-Domain: AndroidWorld
Evaluate Pass@1/Pass@2/Pass@3 on the 116 AndroidWorld tasks. Training explored, generated tasks, rolled out, and trained only within these 20 apps' ecosystem.
Out-of-Domain: MobileWorld GUI-Only
Tested on the 117-task split with no MobileWorld rollouts, tasks, or feedback used during training. This is the strongest evidence that the adaptation isn't "memorizing the training apps."
In practice, ForgeOwl-8B reaches 41.0% on MobileWorld, validating the method's generalization; but ForgeQwen3-8B (built on general-purpose Qwen3-VL-8B) only moves from 7.6% to 10.3% — cross-domain generalization depends heavily on the base model's own mobile GUI capability.
Verifying the Results
After one round of MobileForge adaptation, check these metrics:
- AndroidWorld Pass@3: should improve by 10+ percentage points over the baseline
- Single-pass success rate on hard tasks: on the GUI-Owl-1.5-8B baseline, hard tasks went 19.3% → 29.8% as a reference point
- Tag-wise failure rate drops: MobileForge improves markedly on capabilities strongly tied to app grounding — verification, search, complex UI, screen reading, repetition, information retrieval; game-playing, multi-app, memorization, and math-counting remain hard
FAQ
- Can't run the Critic without Gemini 2.5 Pro? The paper's ablations show that swapping the Critic's decision model to Qwen3-VL-8B still lifts baseline Pass@1 from 40.5% to 44.8%. The closed loop doesn't force dependence on any particular closed-source evaluator.
- Should failed tasks be thrown away? No. MobileForge's best policy is to keep all-failure and partial-success tasks and use step-level feedback to recover the reasonable local actions inside them.
- Are tasks generated from landing screens enough? No. Take Broccoli as an example: basing tasks only on landing screens concentrates 27.3% of them on home-page features like recipe deletion; a curriculum based on exploration trajectories covers much broader functionality — shopping lists, cooking assistant, meal planning, settings, and more.
Primary sources: