Harness Engineering: Rocketing AI Coding Success Rates from 20% to 100%

·Toolin Editorial Team

A conclusion verified by Anthropic and OpenAI at the same time: when AI coding agents fail, the fault lies not in the model but in the Harness. Five steps to build your first Harness setup

Harness Engineering: Rocketing AI Coding Success Rates from 20% to 100%

Does your AI coding agent regularly "confidently deliver a pile of code that doesn't run"? The problem most likely isn't the model itself — it's whether you've fitted it with a Harness.

Anthropic and OpenAI, almost simultaneously in 2026, validated the same conclusion with experiments: when AI coding agents keep failing, the problem isn't the model — it's the Harness infrastructure outside the model. The same Opus 4.5 model, run bare, burned $9 and failed every time; fitted with a Harness, it spent $200 and hit a 100% success rate.

What Is a Harness

A Harness isn't a tool, and it isn't a prompting trick — it's the complete engineering infrastructure built around an AI coding agent, made up of five subsystems:

The five Harness subsystems

SubsystemWhat it solvesCorresponding files
InstructionsThe agent doesn't know project conventions and writes code blindAGENTS.md / CLAUDE.md
ToolsOut-of-bounds operations: rm -rf, git push --forcesettings.json / config.toml
EnvironmentRuns fine on this machine, dead on arrival in CIsetup.sh / Dockerfile
StateCross-session amnesia, writes conflicting codePROGRESS.md
FeedbackDeclares victory early, the code doesn't even runtype check / test / lint

Two Controlled Experiments

The Anthropic Experiment

The same Opus 4.5 model, the same coding problem:

  • Run bare: $9 spent, 0% success — messy code style, destructive commands, no tests run
  • With a Harness: $200 spent, 100% success — the extra $191 all went into the verification loop

Anthropic controlled experiment data

The OpenAI Experiment

The Codex team validated this on real repositories with millions of lines. The experiment changed exactly one thing -- an AGENTS.md file, under 100 lines of markdown, added at the repository root.

The OpenAI Codex experiment

Build Your Harness in Five Steps

The five steps below can all be done in a text editor, adding up to no more than 200 lines of configuration.

Step 1: Create AGENTS.md (or CLAUDE.md)

Create a markdown file in the repository root. The OpenAI camp calls it AGENTS.md; the Anthropic camp calls it CLAUDE.md. Codex, Claude Code, and Cursor read it automatically at startup and inject it into the system prompt.

Write at least three blocks of content:

# Project
This is an e-commerce admin backend built with Next.js + Prisma

# Forbidden
- Never run git push --force
- Never delete the migrations directory
- Never use npm (use pnpm)

# Done means
- pnpm typecheck passes
- pnpm test all green
- pnpm lint zero errors

In under 15 lines, project conventions go from something you repeat over and over to something auto-injected at startup.

AGENTS.md configuration example

Step 2: Configure Permissions

Restrict which commands the agent can invoke.

Claude Code uses .claude/settings.json; Codex uses ~/.codex/config.toml.

{
  "permissions": {
    "allow": ["pnpm install", "pnpm test", "pnpm typecheck"],
    "deny": ["rm -rf", "git push --force", "DROP TABLE"]
  }
}

Allowed commands run directly, denied commands are refused outright, and gray areas pop a confirmation.

Permissions configuration example

Step 3: Write setup.sh to Lock the Environment

Lock down dependency versions and runtime configuration. If you already have a Dockerfile / devcontainer.json you can skip this; otherwise write a setup.sh.

The key line:

pnpm install --frozen-lockfile

--frozen-lockfile ensures the agent can't upgrade any dependency on its own.

Environment lock configuration

Step 4: Create PROGRESS.md

touch PROGRESS.md

Four sections: completed, in progress, to do, known issues. Commit it to git and maintain it as part of the project itself.

Codify the convention in AGENTS.md:

## Rules
- First thing in a new session: read PROGRESS.md
- On task completion or checkpoint change: write back to PROGRESS.md immediately
- On conflict, code wins -- the repo is the single source of truth

PROGRESS.md example

Step 5: Codify the Definition of Done (Most Critical)

State the verification commands at the end of AGENTS.md:

## Done Definition
Task is NOT done until ALL of these pass:
- pnpm typecheck (exit code 0)
- pnpm test (exit code 0)
- pnpm lint (exit code 0)
- pnpm build (exit code 0)

If an exit code isn't 0, the task isn't done. If your project doesn't have these commands yet, set them up today.

Core lesson: get the first four steps right but skip the fifth, and it's still all wasted. Without a feedback loop, a Harness might as well not be installed.

Three Deadly Failure Modes

The Anthropic and OpenAI experiments point to the three most common agent failures:

1. Declaring Victory Too Early

The agent finishes 500 lines of features and outputs "done." Merge the code -- CI goes red, type check reports 12 errors, and not a single unit test has run.

Fix: the feedback subsystem. Hand the verdict to exit codes -- exit code != 0, task != done.

2. Context Anxiety

A long task hits 70%, and context Tokens are nearly full. The agent starts rushing -- skipping tests, deleting edge-case handling, wrapping things up with stubs.

Fix: the state subsystem + proactive restarts. When context Token usage passes 70%, deliberately stop, finish writing the checkpoint, and start a new session.

3. Cross-Session Amnesia

The first session writes the user module; the second session writes getUserById all over again, and the interface signatures conflict.

Fix: PROGRESS.md maintains the completed-features list + AGENTS.md states the read-it-first convention.

The three failure modes

The Core Conclusion

Model capability sets the ceiling; the Harness decides how much of that ceiling you actually reach.

Without a Harness, Opus 4.5 produces code that can't even compile; with one, even a tier-smaller model delivers reliably. Rather than waiting for the next stronger model, install the Harness first.

Resources:

Related articles

Running LLMs on-Device, Hands-On: Building a Local Agentic Coding Environment with Pi + LM Studio
AI Tutorials

Running LLMs on-Device, Hands-On: Building a Local Agentic Coding Environment with Pi + LM Studio

With the Pi agent framework, an LM Studio inference server, and a Docker sandbox, you can run the Gemma 4 series on a local M2 Mac for agentic coding, linting, and unit tests — reaching roughly 75% of frontier-model accuracy.

Toolin Editorial Team
WeChat "Xiaowei" in Gray-Scale Testing, Hands-On: 12 Entry Points That Tuck AI into Your Information Feed
AI Products

WeChat "Xiaowei" in Gray-Scale Testing, Hands-On: 12 Entry Points That Tuck AI into Your Information Feed

A hands-on look at WeChat's native AI assistant "Xiaowei" in gray-scale testing. The main model is WeLM, and 12 entry points cover high-frequency scenarios like chat history retrieval, document summaries, Official Account digests, and local life services.

Toolin Editorial Team
Baidu DuMate, A Practical Guide: From Installation to Office Automation
AI Tutorials

Baidu DuMate, A Practical Guide: From Installation to Office Automation

A full walkthrough of DuMate, the general-purpose office agent from Baidu, covering installation, skills, app connections, and automation — get up and running in 3 minutes and hand your daily office chores to AI.

Toolin Editorial Team
DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories
AI Products

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set, Teaching Code Agents to Build Repositories

Renmin University's Gaoling School has released the DeNovoSWE dataset — 4818 real task instances that train Code Agents to generate complete repositories from documentation, lifting Qwen3-30B from 5.8% to 47.2% on BeyondSWE-Doc2Repo.

Toolin Editorial Team
Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises
AI Products

Doubao Seed 2.1 Pro, Hands-On: Coding Enters the Top Tier, with Multimodal Surprises

A hands-on review of ByteDance's Doubao Seed 2.1 Pro: agent coding and multimodal capability have crossed the production-ready line, including rebuilding front-end interactions from screenshots, at a price nearly 80% lower than Claude Opus 4.6.

Toolin Editorial Team
Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation
AI Products

Hyper3D Rodin Gen-2.5: A Million Polygons in 4 Seconds as Thinking Comes to 3D Generation

Deemos has released Hyper3D Rodin Gen-2.5, the first to bring an LLM-like Thinking mechanism to 3D generation — million-polygon models in 4 seconds, 10-million-polygon precision, and native 12K texturing.

Toolin Editorial Team