PawBench: An Open-Source Agent Benchmark with 4050 Test Cells to Help You Pick Models and Frameworks

·Toolin Editorial Team

PawBench v1.0 builds an agent evaluation set of 150 tasks and 4050 test cells, putting base models and runtime harnesses into the same system to help you find the best model + harness combination.

PawBench: An Open-Source Agent Benchmark with 4050 Test Cells to Help You Pick Models and Frameworks

When an agent task fails, is it that the model "didn't think it through," or that the tools and environment "weren't set up right"? That question has always been hard to answer. PawBench v1.0 puts the base model and the runtime harness into a single evaluation system, helping you find the best combination.

What Is PawBench

PawBench is an open-source evaluation benchmark for personal-assistant and general-agent scenarios. It doesn't just build a model leaderboard; it cross-evaluates "model + harness + task" as a trio.

The evaluation matrix: 9 models x 3 harnesses x 150 tasks = 4,050 test cells

The three harnesses are Hermes, OpenClaw, and QwenPaw. All tasks run in a Docker sandbox, with execution traces and environment snapshots preserved.

PawBench evaluation architecture

A Five-Dimension Label System for the 150 Tasks

Every task is labeled along 5 dimensions:

  1. Use case: office collaboration, software engineering, automation scripting, multimodal content generation
  2. Atomic capability: tool calling, skill usage, planning, logical reasoning, self-verification
  3. Complexity: L1 / L2 / L3
  4. Input modality: text-only vs. multimodal
  5. Runtime environment: offline sandbox vs. requires internet

Key Findings

Finding 1: The Harness Can Swing Model Performance

Swap nothing but the harness on the same model and the score gap can reach 11.5 points. The most telling example is qwen3.6-35b-a3b: it scored 70.4 under QwenPaw but only 68.2 under Hermes.

Causes include:

  • No artifact-level hard verification: the harness takes the model's word for "I'm done" without checking whether files actually landed
  • Poor path awareness: nobody tells the model what its current working directory is, so it writes to the wrong location
  • Tool count overload: Hermes carries about 65 tools, OpenClaw about 30, QwenPaw about 15. Too many tools crowd the context and add decision burden

Finding 2: Proactive Skill Discovery Is a Weak Spot

All three harnesses performed poorly on the 17 skill tasks. The core problem: harnesses never proactively scan the workspace for skill files and rely only on the globally pre-installed skill list.

Finding 3: Web Search Depends on Default Availability

Hermes's search tool needs an API key configured before it turns on — at zero configuration it is "locked out." OpenClaw, by contrast, supports keyless DuckDuckGo search that works straight out of the box.

Four Principles of Harness Design

Based on the results, PawBench offers 4 directly actionable design principles:

PrincipleKey point
Inform FullyTell the model explicitly about its runtime environment: cwd, workspace, output directories, available resources
Equip on DemandCritical tools available by default; tool count matched to the model's attention budget
Monitor ActivelyCheck whether artifacts actually landed instead of taking the model's "done" at face value
Recover GracefullyOn anomalies, inject current state, explain what's missing, and give one chance to self-correct

Who It's For

  • Agent users: helps you pick a better model + harness combination
  • Harness developers: a 4050-cell control matrix plus slice analysis for side-by-side self-checks, failure profiling, and regression validation
  • Researchers: analyze agent capabilities along different dimensions using the five-dimension label system

How to Get It

The project is open source — just search agentscope-ai/PawBench on GitHub. It supports plugging in new harnesses, submitting evaluation results for new models, and contributing new tasks.

Related articles

Xiaomi MiMo UltraSpeed: A Trillion-Parameter Model Running at 1000 tokens/s
AI Products

Xiaomi MiMo UltraSpeed: A Trillion-Parameter Model Running at 1000 tokens/s

Xiaomi's MiMo-V2.5-Pro UltraSpeed delivers 1000 tokens/s on a trillion-parameter model running on generic 8-GPU hardware — and it's the full-strength Pro version, not a dumbed-down Flash variant. The API is live and taking applications.

Toolin Editorial Team
Tabbit: A Permanently Free AI Browser with 10+ Top Models at Your Fingertips
AI Products

Tabbit: A Permanently Free AI Browser with 10+ Top Models at Your Fingertips

Meituan launches AI browser Tabbit V1.0 with core features permanently free, 10+ top domestic large models and agent capabilities built in, plus one-click access to 300+ ready-made trick skills.

Toolin Editorial Team
Finding Overseas Influencers with AI Employees: A Hands-On AhaCreator Guide
AI Tutorials

Finding Overseas Influencers with AI Employees: A Hands-On AhaCreator Guide

A step-by-step walkthrough of the full overseas influencer marketing workflow with AhaCreator — from creator sourcing and content review to cross-border payouts. Ideal for indie developers and teams going global.

Toolin Editorial Team
AI Long-Video Generation: A Head-to-Head Review of Two Open-Source Frameworks
AI Products

AI Long-Video Generation: A Head-to-Head Review of Two Open-Source Frameworks

VideoClaw and JoyAI-Echo, two open-source frameworks released the same day, tackle AI long-video consistency through multi-agent collaboration and cross-modal memory banks respectively — this article compares their technical approaches.

Toolin Editorial Team
ChatGPT's Memory System Gets a Full Overhaul: Dreaming V3 Is Live
AI Products

ChatGPT's Memory System Gets a Full Overhaul: Dreaming V3 Is Live

OpenAI ships the new Dreaming V3 memory architecture — ChatGPT now "dreams" in the background to organize what it knows about you. Free access for 1 billion users for the first time, with doubled memory capacity for Plus/Pro.

Toolin Editorial Team
Cloudflare Integrates Claude Managed Agents: A Developer's Practical Guide
AI Products

Cloudflare Integrates Claude Managed Agents: A Developer's Practical Guide

Cloudflare adds support for Claude Managed Agents, letting developers run Claude agents on Cloudflare's platform, connect to private systems, and deploy AI agents securely.

Toolin Editorial Team