Harness Engineering: Rocketing AI Coding Success Rates from 20% to 100%
A conclusion verified by Anthropic and OpenAI at the same time: when AI coding agents fail, the fault lies not in the model but in the Harness. Five steps to build your first Harness setup


Harness Engineering: Rocketing AI Coding Success Rates from 20% to 100%
A conclusion verified by Anthropic and OpenAI at the same time: when AI coding agents fail, the fault lies not in the model but in the Harness. Five steps to build your first Harness setup
Does your AI coding agent regularly "confidently deliver a pile of code that doesn't run"? The problem most likely isn't the model itself — it's whether you've fitted it with a Harness.
Anthropic and OpenAI, almost simultaneously in 2026, validated the same conclusion with experiments: when AI coding agents keep failing, the problem isn't the model — it's the Harness infrastructure outside the model. The same Opus 4.5 model, run bare, burned $9 and failed every time; fitted with a Harness, it spent $200 and hit a 100% success rate.
What Is a Harness
A Harness isn't a tool, and it isn't a prompting trick — it's the complete engineering infrastructure built around an AI coding agent, made up of five subsystems:

| Subsystem | What it solves | Corresponding files |
|---|---|---|
| Instructions | The agent doesn't know project conventions and writes code blind | AGENTS.md / CLAUDE.md |
| Tools | Out-of-bounds operations: rm -rf, git push --force | settings.json / config.toml |
| Environment | Runs fine on this machine, dead on arrival in CI | setup.sh / Dockerfile |
| State | Cross-session amnesia, writes conflicting code | PROGRESS.md |
| Feedback | Declares victory early, the code doesn't even run | type check / test / lint |
Two Controlled Experiments
The Anthropic Experiment
The same Opus 4.5 model, the same coding problem:
- Run bare: $9 spent, 0% success — messy code style, destructive commands, no tests run
- With a Harness: $200 spent, 100% success — the extra $191 all went into the verification loop

The OpenAI Experiment
The Codex team validated this on real repositories with millions of lines. The experiment changed exactly one thing -- an AGENTS.md file, under 100 lines of markdown, added at the repository root.

Build Your Harness in Five Steps
The five steps below can all be done in a text editor, adding up to no more than 200 lines of configuration.
Step 1: Create AGENTS.md (or CLAUDE.md)
Create a markdown file in the repository root. The OpenAI camp calls it AGENTS.md; the Anthropic camp calls it CLAUDE.md. Codex, Claude Code, and Cursor read it automatically at startup and inject it into the system prompt.
Write at least three blocks of content:
# Project
This is an e-commerce admin backend built with Next.js + Prisma
# Forbidden
- Never run git push --force
- Never delete the migrations directory
- Never use npm (use pnpm)
# Done means
- pnpm typecheck passes
- pnpm test all green
- pnpm lint zero errorsIn under 15 lines, project conventions go from something you repeat over and over to something auto-injected at startup.

Step 2: Configure Permissions
Restrict which commands the agent can invoke.
Claude Code uses .claude/settings.json; Codex uses ~/.codex/config.toml.
{
"permissions": {
"allow": ["pnpm install", "pnpm test", "pnpm typecheck"],
"deny": ["rm -rf", "git push --force", "DROP TABLE"]
}
}Allowed commands run directly, denied commands are refused outright, and gray areas pop a confirmation.

Step 3: Write setup.sh to Lock the Environment
Lock down dependency versions and runtime configuration. If you already have a Dockerfile / devcontainer.json you can skip this; otherwise write a setup.sh.
The key line:
pnpm install --frozen-lockfile--frozen-lockfile ensures the agent can't upgrade any dependency on its own.

Step 4: Create PROGRESS.md
touch PROGRESS.mdFour sections: completed, in progress, to do, known issues. Commit it to git and maintain it as part of the project itself.
Codify the convention in AGENTS.md:
## Rules
- First thing in a new session: read PROGRESS.md
- On task completion or checkpoint change: write back to PROGRESS.md immediately
- On conflict, code wins -- the repo is the single source of truth
Step 5: Codify the Definition of Done (Most Critical)
State the verification commands at the end of AGENTS.md:
## Done Definition
Task is NOT done until ALL of these pass:
- pnpm typecheck (exit code 0)
- pnpm test (exit code 0)
- pnpm lint (exit code 0)
- pnpm build (exit code 0)If an exit code isn't 0, the task isn't done. If your project doesn't have these commands yet, set them up today.
Core lesson: get the first four steps right but skip the fifth, and it's still all wasted. Without a feedback loop, a Harness might as well not be installed.
Three Deadly Failure Modes
The Anthropic and OpenAI experiments point to the three most common agent failures:
1. Declaring Victory Too Early
The agent finishes 500 lines of features and outputs "done." Merge the code -- CI goes red, type check reports 12 errors, and not a single unit test has run.
Fix: the feedback subsystem. Hand the verdict to exit codes -- exit code != 0, task != done.
2. Context Anxiety
A long task hits 70%, and context Tokens are nearly full. The agent starts rushing -- skipping tests, deleting edge-case handling, wrapping things up with stubs.
Fix: the state subsystem + proactive restarts. When context Token usage passes 70%, deliberately stop, finish writing the checkpoint, and start a new session.
3. Cross-Session Amnesia
The first session writes the user module; the second session writes getUserById all over again, and the interface signatures conflict.
Fix: PROGRESS.md maintains the completed-features list + AGENTS.md states the read-it-first convention.

The Core Conclusion
Model capability sets the ceiling; the Harness decides how much of that ceiling you actually reach.
Without a Harness, Opus 4.5 produces code that can't even compile; with one, even a tier-smaller model delivers reliably. Rather than waiting for the next stronger model, install the Harness first.
Resources:
Toolin Editorial Team
Categories
Related articles

Running LLMs on-Device, Hands-On: Building a Local Agentic Coding Environment with Pi + LM Studio
With the Pi agent framework, an LM Studio inference server, and a Docker sandbox, you can run the Gemma 4 series on a local M2 Mac for agentic coding, linting, and unit tests — reaching roughly 75% of frontier-model accuracy.

WeChat "Xiaowei" in Gray-Scale Testing, Hands-On: 12 Entry Points That Tuck AI into Your Information Feed
A hands-on look at WeChat's native AI assistant "Xiaowei" in gray-scale testing. The main model is WeLM, and 12 entry points cover high-frequency scenarios like chat history retrieval, document summaries, Official Account digests, and local life services.

Baidu DuMate, A Practical Guide: From Installation to Office Automation
A full walkthrough of DuMate, the general-purpose office agent from Baidu, covering installation, skills, app connections, and automation — get up and running in 3 minutes and hand your daily office chores to AI.

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
Renmin University's Gaoling School has released the DeNovoSWE dataset — 4818 real task instances that train Code Agents to generate complete repositories from documentation, lifting Qwen3-30B from 5.8% to 47.2% on BeyondSWE-Doc2Repo.

Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises
A hands-on review of ByteDance's Doubao Seed 2.1 Pro: agent coding and multimodal capability have crossed the production-ready line, including rebuilding front-end interactions from screenshots, at a price nearly 80% lower than Claude Opus 4.6.

Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation
Deemos has released Hyper3D Rodin Gen-2.5, the first to bring an LLM-like Thinking mechanism to 3D generation — million-polygon models in 4 seconds, 10-million-polygon precision, and native 12K texturing.