PawBench: An Open-Source Agent Benchmark with 4050 Test Cells to Help You Pick Models and Frameworks
PawBench v1.0 builds an agent evaluation set of 150 tasks and 4050 test cells, putting base models and runtime harnesses into the same system to help you find the best model + harness combination.


PawBench: An Open-Source Agent Benchmark with 4050 Test Cells to Help You Pick Models and Frameworks
PawBench v1.0 builds an agent evaluation set of 150 tasks and 4050 test cells, putting base models and runtime harnesses into the same system to help you find the best model + harness combination.
When an agent task fails, is it that the model "didn't think it through," or that the tools and environment "weren't set up right"? That question has always been hard to answer. PawBench v1.0 puts the base model and the runtime harness into a single evaluation system, helping you find the best combination.
- Project: https://github.com/agentscope-ai/PawBench
- Leaderboard: https://agentscope-ai.github.io/PawBench
What Is PawBench
PawBench is an open-source evaluation benchmark for personal-assistant and general-agent scenarios. It doesn't just build a model leaderboard; it cross-evaluates "model + harness + task" as a trio.
The evaluation matrix: 9 models x 3 harnesses x 150 tasks = 4,050 test cells
The three harnesses are Hermes, OpenClaw, and QwenPaw. All tasks run in a Docker sandbox, with execution traces and environment snapshots preserved.

A Five-Dimension Label System for the 150 Tasks
Every task is labeled along 5 dimensions:
- Use case: office collaboration, software engineering, automation scripting, multimodal content generation
- Atomic capability: tool calling, skill usage, planning, logical reasoning, self-verification
- Complexity: L1 / L2 / L3
- Input modality: text-only vs. multimodal
- Runtime environment: offline sandbox vs. requires internet
Key Findings
Finding 1: The Harness Can Swing Model Performance
Swap nothing but the harness on the same model and the score gap can reach 11.5 points. The most telling example is qwen3.6-35b-a3b: it scored 70.4 under QwenPaw but only 68.2 under Hermes.
Causes include:
- No artifact-level hard verification: the harness takes the model's word for "I'm done" without checking whether files actually landed
- Poor path awareness: nobody tells the model what its current working directory is, so it writes to the wrong location
- Tool count overload: Hermes carries about 65 tools, OpenClaw about 30, QwenPaw about 15. Too many tools crowd the context and add decision burden
Finding 2: Proactive Skill Discovery Is a Weak Spot
All three harnesses performed poorly on the 17 skill tasks. The core problem: harnesses never proactively scan the workspace for skill files and rely only on the globally pre-installed skill list.
Finding 3: Web Search Depends on Default Availability
Hermes's search tool needs an API key configured before it turns on — at zero configuration it is "locked out." OpenClaw, by contrast, supports keyless DuckDuckGo search that works straight out of the box.
Four Principles of Harness Design
Based on the results, PawBench offers 4 directly actionable design principles:
| Principle | Key point |
|---|---|
| Inform Fully | Tell the model explicitly about its runtime environment: cwd, workspace, output directories, available resources |
| Equip on Demand | Critical tools available by default; tool count matched to the model's attention budget |
| Monitor Actively | Check whether artifacts actually landed instead of taking the model's "done" at face value |
| Recover Gracefully | On anomalies, inject current state, explain what's missing, and give one chance to self-correct |
Who It's For
- Agent users: helps you pick a better model + harness combination
- Harness developers: a 4050-cell control matrix plus slice analysis for side-by-side self-checks, failure profiling, and regression validation
- Researchers: analyze agent capabilities along different dimensions using the five-dimension label system
How to Get It
The project is open source — just search agentscope-ai/PawBench on GitHub. It supports plugging in new harnesses, submitting evaluation results for new models, and contributing new tasks.
Toolin Editorial Team
Categories
Related articles

Xiaomi MiMo UltraSpeed: A Trillion-Parameter Model Running at 1000 tokens/s
Xiaomi's MiMo-V2.5-Pro UltraSpeed delivers 1000 tokens/s on a trillion-parameter model running on generic 8-GPU hardware — and it's the full-strength Pro version, not a dumbed-down Flash variant. The API is live and taking applications.

Tabbit: A Permanently Free AI Browser with 10+ Top Models at Your Fingertips
Meituan launches AI browser Tabbit V1.0 with core features permanently free, 10+ top domestic large models and agent capabilities built in, plus one-click access to 300+ ready-made trick skills.

Finding Overseas Influencers with AI Employees: A Hands-On AhaCreator Guide
A step-by-step walkthrough of the full overseas influencer marketing workflow with AhaCreator — from creator sourcing and content review to cross-border payouts. Ideal for indie developers and teams going global.

AI Long-Video Generation: A Head-to-Head Review of Two Open-Source Frameworks
VideoClaw and JoyAI-Echo, two open-source frameworks released the same day, tackle AI long-video consistency through multi-agent collaboration and cross-modal memory banks respectively — this article compares their technical approaches.

ChatGPT's Memory System Gets a Full Overhaul: Dreaming V3 Is Live
OpenAI ships the new Dreaming V3 memory architecture — ChatGPT now "dreams" in the background to organize what it knows about you. Free access for 1 billion users for the first time, with doubled memory capacity for Plus/Pro.

Cloudflare Integrates Claude Managed Agents: A Developer's Practical Guide
Cloudflare adds support for Claude Managed Agents, letting developers run Claude agents on Cloudflare's platform, connect to private systems, and deploy AI agents securely.