Bun Rewrote Its Core from Zig to Rust with Claude: A Million-Line Agentic Engineering Effort

·Toolin Editorial Team

Bun founder Jarred Sumner used Claude Code to rewrite Bun's core from Zig to Rust — roughly 1 million lines of code in about 10 days. This piece breaks down the engineering significance of the agentic rewrite, what is verifiable, and where its applicability ends.

Bun Rewrote Its Core from Zig to Rust with Claude: A Million-Line Agentic Engineering Effort

In mid-July, Bun founder Jarred Sumner completed an unusual engineering experiment: using Claude Code (Fable 5) to rewrite Bun's core from Zig to Rust — roughly 1 million lines of code, taking about 9 to 11 days end to end. It drew wide discussion not simply because "AI wrote the code," but because it pushed agentic coding to a new scale — not writing a function or a component, but rewriting the internals of a widely used runtime. This article isn't reheating the "AI replaces programmers" take; it treats this as an agentic coding engineering case study: what's verifiable, what it means for tool selection, and which signals deserve skepticism.

Note: This article takes an engineering-case / showcase perspective. If you're after a how-to tutorial on the Claude Code agentic loop itself, see this site's 2026-07-16 article, "Claude Code Agentic Loops."

What This Rewrite Actually Was

First, pin down the factual layer to avoid being misled by headlines:

DimensionVerifiable information
Who did itBun founder Jarred Sumner
Tool usedClaude Code (model: Fable 5)
TaskRewrite Bun's core from Zig to Rust
Code sizeRoughly 1 million lines
DurationAbout 9 to 11 days
CostHeadline retellings say "about $160,000"; not fully confirmed by independent sources

The "$160,000" figure needs a special note: it is a retelling from some media headlines and has not been fully confirmed by independent sources. There is also another figure often cited in the community — a $570K enterprise agentic coding cost case mentioned by a CIO — which is a different event entirely; don't conflate the two. In your own decision-making, treat "cost" as a variable awaiting verification, not a conclusion.

Why This Episode Is Worth Attention

Most agentic coding cases over the past two years stayed at the scale of "write a demo" or "refactor a module." Bun's rewrite keeps getting discussed because it pushed agentic coding along three new dimensions:

  • Scale: a codebase in the millions of lines
  • Risk: Bun is a production-grade runtime used by a large developer base
  • Language migration: not a same-language rewrite but a cross-language Zig → Rust migration

Cross-language migration is far harder than same-language refactoring — it means remapping the language runtime, memory model, and error-handling paradigms, not just swapping syntax. That makes it exactly the "stress-test scenario" for agentic coding: if Claude Code can make steady progress on a task like this, its capability boundary on ordinary engineering work deserves re-evaluation.

Background: Bun's Community Context

Two contextual signals are unavoidable in understanding this rewrite:

  1. Bun has reportedly been acquired by Anthropic: this means the agentic rewrite wasn't an ordinary third party using Claude — it happened inside the close relationship between acquirer and acquired.
  2. Community controversy: Zig's author publicly questioned the quality of AI-generated code. This reflects a broader concern among language/toolchain maintainers about AI-generated code — can it be maintained, does it follow the language's idioms, will it become a long-term burden?

These two signals mean you can't simply extrapolate this case to ordinary teams. Jarred's own familiarity with the Bun codebase, the close relationship with Anthropic, and deeply tuned Claude Code workflows are conditions ordinary users can't easily reproduce.

How to Understand the "Agentic Rewrite" Paradigm

From this case we can abstract several features that distinguish an agentic rewrite from traditional refactoring:

[Traditional refactoring]
Human engineers → rewrite file by file → run tests → fix → next file
   Low parallelism, steady cadence, auditable

[Agentic rewrite]
Human sets goals and constraints → agent rewrites in bulk and self-checks
   → human reviews key diffs → agent keeps pushing forward
   Key: humans stay in the loop for "decide + verify," agents do "execute + self-check"

This paradigm actually demands more from its users, not less:

  • You have to define clear acceptance criteria (otherwise the agent doesn't know when to stop)
  • You have to be able to read large-scale diffs (otherwise you can't review)
  • You have to set cost and risk ceilings (otherwise bills and errors spiral together)

This is exactly what the 2026-07-16 article "Claude Code Agentic Loops" emphasized — every effective loop must explicitly design its goal and stop conditions. Bun's case is, at its core, a Goal-driven Loop at extraordinary scale.

What This Case Teaches About Tool Selection

If you're evaluating whether "large-scale refactoring with agentic coding" is feasible, this case yields several judgment dimensions:

  • Testability of the task: a runtime like Bun has a complete test suite and benchmarks, so agent output can be verified automatically. If your project's test coverage is thin, the risk of an agentic rewrite compounds exponentially.
  • Depth of human domain knowledge: Jarred is Bun's creator and knows the story behind every design decision. For an ordinary person to use Claude on a migration of the same scale, the prerequisite is comparable depth of knowledge of your own codebase.
  • Toolchain maturity: cross-language migration depends on how sound the target language's (Rust) toolchain, linting, and formatters are — that determines whether agent-generated code can be verified quickly.
  • Cost controllability: the cost of a long agentic loop is token consumption × rounds. A migration at the million-line scale amplifies token costs fast. Run a small-sample trial before starting to estimate full-scale cost.

Where It Fits and Where It Doesn't

Scenarios That Can Learn from This Case

  • Language migration or framework upgrades in large codebases: complete tests, human familiarity with the codebase, capacity to absorb long-loop costs
  • Cross-language rewrite prototypes: have the agent produce a runnable Rust/Go/TypeScript version first, then polish by hand
  • Stress-testing agentic coding capability: use it as a benchmark to evaluate how stable different models/tools are on long tasks

Scenarios Where It Should Not Be Copied Directly

  • Projects with thin test coverage: with no automated acceptance checks, agent output can't be verified
  • Urgent refactoring on business-critical paths: long agentic loops are uncertain in both time and quality, a poor fit for tight deadlines
  • Open-source projects with strict style and idiom requirements: AI-generated code often violates language idioms and needs heavy manual polishing

Before You Use This

  • "$160,000" is unverified data: don't plug this number directly into your cost estimates. Run a small-sample trial with your own task size and token prices first, get the actual cost per unit of code, then extrapolate.
  • Don't ignore community criticism: the Zig author's objections represent legitimate concerns from language maintainers — maintainability of AI-generated code, idiom conformance, long-term technical debt. Before adopting an agentic rewrite, make sure you are able to review output quality.
  • Don't extrapolate a "founder did it himself" case to ordinary teams: Jarred's depth of domain knowledge, close ties to Anthropic, and deeply tuned tooling are implicit conditions — none of them automatically transfer to ordinary usage.

References

  • Bun official and community discussions (mid-July 2026)
  • Related tutorial on this site: "Claude Code Agentic Loops: 4 Engineering Practices for Moving from Writing Prompts to Designing Loops" (2026-07-16)
  • Anthropic: Building Effective Agents