Alibaba Open-Sources Qwen-AgentWorld: The First Language World Model, One Model Simulating 7 Types of Environments
Alibaba's Qwen team has open-sourced Qwen-AgentWorld, the first native language world model. A single model covers seven agent environments — MCP / Search / Terminal / SWE / Web / OS / Android — with the AgentWorldBench evaluation benchmark released alongside it.


Alibaba Open-Sources Qwen-AgentWorld: The First Language World Model, One Model Simulating 7 Types of Environments
Alibaba's Qwen team has open-sourced Qwen-AgentWorld, the first native language world model. A single model covers seven agent environments — MCP / Search / Terminal / SWE / Web / OS / Android — with the AgentWorldBench evaluation benchmark released alongside it.
Alibaba's Qwen team has released the first native language world model (LWM), Qwen-AgentWorld, built specifically for developing and training AI agents. Its core capability: letting an agent simulate environment feedback internally before acting, then decide. A single model covers 7 types of agent interaction environments, and the AgentWorldBench evaluation benchmark ships alongside it — both the model weights and the benchmark are open-sourced.
This piece breaks down what Qwen-AgentWorld is, how it works, how it performs, and how to get it.
What it is
Qwen-AgentWorld is a language world model available in two parameter sizes:
- Qwen-AgentWorld-35B-A3B (open weights)
- Qwen-AgentWorld-397B-A17B
It is not meant to replace interaction with real environments — real environments remain the gold standard for ensuring agent behavior is reliable. What the LWM provides is a complementary path: scalability and controllability beyond what real environments offer, plus an internalized capacity for world prediction.
The figure above shows a simulated phone system: the left side is the phone's initial state, and the right side is the predicted outcome of the agent tapping the delete icon in the toolbar.
Two core highlights
Highlight 1: environment modeling as a training objective from pretraining onward
Environment modeling runs through the entire CPT → SFT → RL pipeline. Previously, training general-purpose foundation models usually meant teaching AI to understand environments and predict action outcomes only after training was over. Qwen-AgentWorld moves this step forward into pretraining, having been trained on more than 10 million real environment interaction trajectories.

Highlight 2: a single model covering 7 types of environments
| Category | Environments covered |
|---|---|
| Text-based | MCP, Search, Terminal, SWE |
| GUI-based | Web, OS, Android |
Environment observations in the three GUI domains are represented as renderable code (accessibility tree XML, HTML, UI hierarchy markup) rather than pixel frames, which allows purely text-based world modeling to cover visual environments and enables cross-domain knowledge transfer.

It can simulate desktop systems (such as predicting the outcome of clicking "File" > "Print" in the menu bar) and website interactions (such as predicting the outcome of clicking an "Add User" button).
Performance: overall simulation quality beats Claude Opus 4.8
The accompanying AgentWorldBench was built from real environment interaction observations of 5 frontier models across 9 established evaluation sets, scored along five dimensions: format, factual accuracy, consistency, realism, and quality.
- Qwen-AgentWorld-397B-A17B achieved the highest overall average score on AgentWorldBench at 58.71, surpassing GPT-5.4 (58.25), Claude Opus 4.8, and Gemini 3.1 Pro. Its lead is most pronounced in the Terminal and SWE domains.
- At the 35B-A3B scale, the three-stage training pipeline lifted the overall average score by 8.66 points, outperforming Claude Sonnet 4.6.

Three emergent reasoning modes
The researchers analyzed 129 chains of thought across the 4 text-based domains and found 3 emergent reasoning patterns:
- Self-correction: the model uses "Wait!" as a trigger for self-correction, revising intermediate predictions (10.4 times per turn on average), covering factual errors, knowledge boundaries, and perspective shifts.
- Information leakage protection: in the Search domain, the model knows the reference answer the agent is searching for, and it prevents leakage by ensuring its summaries do not inadvertently reveal the target.
- Multi-step causal reasoning: for example, when predicting the output of
curl -s localhost:3000 | python3 -m json.tool, it can produce a 6-step chain of reasoning — Node.js missing → server not started → nothing listening on port 3000 → curl fails silently → empty pipe → json.tool throws a JSONDecodeError.
💡 Tip: Qwen-AgentWorld demonstrates that a language world model can serve both as a "decoupled environment simulator" giving agent RL a controllable training environment, and as a "unified agent foundation model" — its pretraining transfers to multi-turn agent tasks spanning seven benchmarks, with no RL fine-tuning for agent tasks whatsoever.
How to get it
Alibaba has open-sourced Qwen-AgentWorld-35B-A3B (model weights) and AgentWorldBench (evaluation benchmark).
- GitHub: https://github.com/QwenLM/Qwen-AgentWorld
- ModelScope: https://modelscope.cn/collections/Qwen/qwen-agentworld
- Hugging Face: https://huggingface.co/collections/Qwen/qwen-agentworld
Who it is for
- Agent R&D teams: treat the LWM as a controllable environment simulator, giving agent RL a training path that scales better than real environments.
- GUI / computer-use agent developers: use Qwen-AgentWorld to predict the outcome of actions on Web / OS / Android, lowering the cost of trial and error.
- Agent foundation model researchers: explore how language world model pretraining transfers to multi-turn agent tasks.