ClawGym: An Open-Source Framework Unifying Agent Training and Evaluation
RUC has open-sourced a full-pipeline Claw Agent framework spanning data, training, and evaluation, with 13.5K executable tasks and support for sandbox-parallel reinforcement learning


ClawGym: An Open-Source Framework Unifying Agent Training and Evaluation
RUC has open-sourced a full-pipeline Claw Agent framework spanning data, training, and evaluation, with 13.5K executable tasks and support for sandbox-parallel reinforcement learning
Large models are moving from "answering questions" to "getting tasks done", but systematic solutions for building data, training models, and evaluating capabilities for Personal Agents (such as OpenClaw-style desktop agents) have been lacking. ClawGym, open-sourced by Renmin University of China and the Zhizhi Research Institute, provides a complete closed loop from data synthesis to training to evaluation, and is currently the most comprehensive OpenClaw training and evaluation resource.
What Is ClawGym
ClawGym is a unified framework for Claw Agents that systematically connects data synthesis, model training, and reliable evaluation. It contains three core modules:
- ClawGym-SynData: The first large-scale synthetic dataset for Claw Agents, containing 13.5K executable tasks
- ClawGym-Agents: Trains Agents on OpenClaw black-box execution trajectories and explores sandbox-parallel reinforcement learning
- ClawGym-Bench: An evaluation benchmark of 200 high-quality tasks covering six categories of workspace scenarios

- GitHub: https://github.com/ClawGym
Why a Dedicated Agent Framework Is Needed
Claw-style environments differ fundamentally from traditional text Q&A, web browsing, or simple tool calls. What the Agent faces is not a static problem but a complex workspace made of files, directories, scripts, spreadsheets, configs, logs, and external tools.
It needs to read files, run commands, analyze data, modify documents, and generate reports across multi-turn interactions, constantly adjusting its actions based on environmental feedback. Every operation changes the workspace state, and subsequent decisions depend on those intermediate states.
Whether a task is complete does not depend on the Agent saying "I'm done", but on whether the final workspace has actually been updated correctly.
This raises four core challenges:
- Tasks are hard to construct: They must cover real workflows and executable operations, not just generate a prompt
- Trajectories are hard to collect: High-quality training trajectories must be reconstructed from black-box execution logs
- Training is hard to stabilize: The reinforcement learning stage requires concurrent rollouts across large numbers of independent sandboxes
- Rewards are hard to define: File, structure, numeric, and multi-dimensional artifact quality all need verification
ClawGym-SynData: 13.5K Executable Tasks
Dual-Route Task Synthesis
To make sure tasks both stay close to real needs and are genuinely executable, ClawGym uses two complementary synthesis routes:
- Persona-driven, top-down: Starts from "what the user wants to do", building user personas, work scenarios, and atomic operation combinations to generate tasks close to real office scenes
- Skill-grounded, bottom-up: Starts from "what the system can do", extracting reusable tool capabilities from OpenClaw skills so tasks land on runnable operations
Automatically Generated Mock Workspaces
Every task comes with an automatically generated lightweight mock workspace (Markdown, JSON, CSV, YAML, config files, and so on) supplying the content that needs to be read, analyzed, and modified during execution.
Hybrid Verification Mechanism
- Code-based verification: Checks objective correctness such as file paths, schemas, numeric computation, and filtering rules
- Rubric-based verification: Assesses subjective qualities such as report clarity, summary faithfulness, and professionalism of expression
ClawGym-Agents: Training From Real Trajectories
ClawGym collects real interaction trajectories through OpenClaw black-box rollouts rather than reimplementing a simplified agent loop. After aggregation, cleaning, and filtering, the trajectories average:
- 13.00 interaction turns
- 18.67K tokens
- 15.82 tool calls
- 3.25 tool types
The Qwen3 model family was put through multi-turn SFT on these trajectories, yielding three models:
| Model | Base | Character |
|---|---|---|
| ClawGym-4B | Qwen3-4B | Lightweight |
| ClawGym-8B | Qwen3-8B | Balanced |
| ClawGym-30B-A3B | Qwen3-30B-A3B | High-performance |
The team also explored sandbox-parallel RL: each task runs in an independent sandbox, with a code verifier providing the outcome reward. Experiments show RL delivers further gains on top of SFT.
ClawGym-Bench: 200 Curated Evaluation Tasks
ClawGym-Bench contains 200 rigorously screened tasks covering six categories of typical workspace scenarios:
- Productivity and collaboration
- Systems and automation
- Analysis and reasoning
- Content and domain support
- Planning and knowledge management
- Software development
Every task passes a dual review of "LLM diagnostic-style checks + human review" to ensure clear instructions, complete resources, and reliable verification.
Experimental Results
Key numbers:
- ClawGym-4B, 8B, and 30B-A3B reach 47.73, 50.24, and 56.82 respectively on ClawGym-Bench, all surpassing their corresponding base models
- ClawGym-30B-A3B surpasses the much larger Qwen3-235B-A23B, showing that high-quality Agent data can make up for model scale
- Trained with ClawGym-SynData alone, the models also post clear gains on the external benchmark PinchBench (ClawGym-30B-A3B reaches 86.00), proving what's learned is not task templates but transferable execution capability
Open-Sourced Resources
The team has open-sourced five core resources:
- ClawGym-Bench evaluation data
- Evaluation code
- ClawGym-Agents model checkpoints
- Training data
- Training code
GitHub: https://github.com/ClawGym
Use Cases
- Agent researchers: Complete training data and an evaluation benchmark
- Model developers: Can use ClawGym-SynData directly for multi-turn SFT and RL
- Agent product teams: Use ClawGym-Bench to assess different models' actual execution capability
- Open-source community: Extend the ClawGym framework with more task types and evaluation dimensions
ClawGym's core value is that it doesn't just ask whether a model can "say the answer", but systematically asks whether the model can complete checkable, verifiable tasks in a workspace. For Personal Agents, this is a key step from conversational ability to execution ability.