Loom: An Engineering State Layer for Coding Agents That Cures Long-Task Local Amnesia

·Toolin Editorial Team

The open-source tool Loom gives coding agents like Claude Code / Codex an independent, structured engineering state layer, enabling long-task auto-checkpointing and zero-cost multi-agent handoff.

Loom: An Engineering State Layer for Coding Agents That Cures Long-Task Local Amnesia

If you've ever fixed bugs with Claude Code or Codex on a real project, you've probably hit this wall: dazzling for the first 3 rounds, collapse by round 10 — the agent flips a bug it already fixed in round 2 back to broken, the session clogs up with errors and junk logs, and you have to manually cut it off and start a new session, then manually narrate the past hour to the new agent in natural language.

Loom exists to solve exactly this "local amnesia." It is a Delivery Harness open-sourced by Valkor together with Zhejiang University's Research Center for Intelligent Computing and Software and UCL's software engineering team, adding an independent, structured engineering state layer for coding agents.

Code repository: github.com/valkor-ai/loom

What Loom Is

A plainer analogy: Loom is a delivery harness that introduces "auto-save points" to a single-player game, decomposing a complex software delivery into a structured state chain that can be resumed at any time.

The core problem it solves is not "how to get the model to write more code" but "how a long task keeps advancing through a real engineering process — verifiably and recoverably — all the way to completion."

When an agent's test run fails, Loom does not dump the failure into the chat context as a piece of terminal text; it captures the failure and structures it as a to-do state that exists independently of the chat history.

Bugs Don't Get Buried

A failure becomes a "hard constraint" on the next action; there is no way for the agent to brush the bug aside in later conversation.

Zero-Cost Multi-Agent Handoff

You can use Claude 4.6 Sonnet to fix logic for the first 5 rounds and switch to GPT-5.5 to run tests in round 6. Once a new agent connects to Loom, reading the structured "delivery state chain" tells it instantly:

  • who it is, where it is, and what was just changed
  • which bug to fix next
  • which files' diffs are finalized and must not be touched

No need to re-read a lengthy chat history — it takes over the game right where it stands.

Why Cranking the Context Window Won't Cure Long Tasks

Blaming long-task failures on a context window that isn't long enough is a false proposition in real engineering logic.

More information does not mean more reliability. Stuffing tens of thousands of lines of compile logs, rounds of diffs, and test output into the context only fills the session with noise; the model gets lost easily, or misjudges locally-runnable code as finished.

The essence of engineering is structure and determinism. Loom's approach is not to make the model see more and remember more, but to help the model filter out the noise and distill only the most essential engineering cues:

  • Which step of the plan are we on?
  • Which unit tests actually pass?
  • Which files' diffs are finalized and off-limits?

Once these key points become structured data that programs can read, the yardstick for AI coding effectiveness shifts from "how many lines of code the model can generate" to "how completely a long task keeps advancing."

Core Features

  • Auto-save points: decomposes the delivery process into a structured, resumable state chain, like a save point in a single-player game.
  • Independent engineering state layer: decoupled from the chat context; key engineering cues exist as structured data that programs can read.
  • Multi-agent / multi-model handoff: a new agent reads the state chain and takes over in place — zero-cost model switching.
  • Failure as constraint: test failures are captured as to-do states and become hard constraints on the next step, never drowned out by chat noise.

Hands-On Experience

Strengths

  • Solves long-task derailment: the agent no longer loses direction as the context stretches; round 10 won't undo the bug fixed in round 2.
  • Unlocks multi-model collaboration: switch to the best-fit model for each task phase — one vendor for logic, another for tests, with no interference.
  • Complements the "crank the context window" route: Loom neither conflicts nor replaces; it is a layer of infrastructure stacked outside existing coding agent workflows.

Costs

  • You need to wire Loom into your existing Claude Code / Codex workflow, which carries some initial setup cost.
  • As a relatively young open-source project, its ecosystem and documentation maturity need to be evaluated on your own.

Use Cases

  • Long-task delivery: 10+ rounds of complex bug fixes, API completion, or feature development within a single session.
  • Multi-model workflows: dynamically switch between Claude, GPT, and other models by task phase without losing intermediate state.
  • Pre-production validation: between a working demo and real software lies a whole layer of trustworthiness verification; Loom provides the state infrastructure.
  • Agent evaluation / fine-tuning data collection: Loom captures agents' dynamic feedback trajectories during execution, providing real engineering corpora for future dynamic benchmarks and fine-tuning.

Positioning and Complementarity

Loom complements rather than overlaps the general agentic loops pattern (loop, feedback, tool calls): loops concern the execution loop within a single task; Loom concerns the persistence of engineering state across tasks, sessions, and models.

Large models learn from static corpora "what perfect final code looks like," but they struggle to learn "how a complex bug actually gets located step by step, fails, compromises, and is finally fixed" — and it is precisely these process trajectories that hold software engineering's most core engineering judgment.

Open-source address: github.com/valkor-ai/loom