Toolin.ai

The Opus 5 "Challenge Loop" Prompt: Making AI Its Own Taskmaster to Grind Out a AAA Game Prototype in 24 Hours

Published · toolin小编

Matt Shumer's "Challenge Loop" prompt drives Opus 5 to iterate on its own with a multi-agent structure of main Agent + builder Agents + judge Agent, paired with Three.js and Blender MCP, replicating an Outer Wilds-caliber browser game solo in 24 hours.

The Opus 5 "Challenge Loop" Prompt: Making AI Its Own Taskmaster to Grind Out a AAA Game Prototype in 24 Hours

An Opus 5 game prompt recently went viral on X: developer Anshu went from zero to a playable browser-based space exploration game, The Long Silence, in 24 hours, solo. The original author, Matt Shumer, used the same prompt to get Opus 5 to produce a Call of Duty-caliber FPS prototype without any external assets.

The core of this prompt is the "Challenge Loop"—it doesn't teach the model how to make games; it installs a referee that won't let the model clock out. This post breaks down its multi-agent structure, the reusable three-step process, and how to port the pattern to your own long-horizon tasks.

Playable build: The Long Silence is open-sourced and playable directly in the browser—longsilence.anshu.dev

What Is the Challenge Loop: Three Agents Applying Pressure to Each Other

Most multi-agent game development pipelines stop at "split tasks → do work." The Challenge Loop adds an independent judge Agent, structured like this:

  • Main Agent: receives the goal, breaks it down itself, and dispatches sub-agents
  • Builder Agents (multiple): produce code, textures, animation, and sound in parallel
  • Judge Agent: runs blind side-by-side comparisons of generated output against real AAA references; anything below the visual bar is rejected and redone

The key is that the Judge Agent is independent—it doesn't report to the Main Agent, so the model can't quietly lower its own grading bar to let itself off the hook. Anything that doesn't match comparable games automatically goes back for rework.

⚠️ Reality check: this prompt cannot literally replicate a AAA title 1:1. Treat it as an extreme-pressure device—one that squeezes out Opus 5's autonomous planning and long-horizon iteration, with you (the user) deciding when to call a stop.

Before You Start

  • Model: Anthropic Claude Opus 5 (run inside Claude Code)
  • Modeling tool: install Blender MCP in Claude Code so the model can call Blender itself
  • Rendering framework: Three.js (60 fps browser target)
  • Budget: 24 hours of model token consumption isn't cheap; estimate your quota in advance
  • GitHub repo: for persisting the skill Opus 5 auto-summarizes

The Three-Step Replication Process

Step 1: Give a baseline goal and let Opus 5 decide the worldview

Give only the minimum hard constraints and let the model decide the rest. For example:

Make a space exploration game with Three.js. The player can walk around and pilot a ship. Visuals shouldn't look too plasticky. Keep it running stably at 60 fps in the browser.

The game's worldview, technical architecture, and art style—all left to Opus 5's own decisions. For assets, once Blender MCP is installed, Claude Code finds the tools and models things itself; you don't hand-feed assets.

💡 Tip: don't over-constrain at this step. Opus 5's long-horizon planning needs room to work; boxing it in early only limits its iteration space.

Step 2: Set a "24-hour continuous goal" and start the Judge Agent

Once the first playable version exists, give Opus 5 a long-window goal demanding across-the-board visual improvement. Multiple sub-agents now start working in parallel:

  • Builder Agents optimize textures, lighting, and animation in parallel
  • The Judge Agent compares screenshots item by item against real AAA space games like Starfield
  • Opus 5 is not allowed to quietly lower the grading bar along the way

Anshu also notes you should give Opus 5 some "carrots" alongside the extreme pressure:

You are capable of reaching this goal—do your best, don't give up.

Check progress from your phone at any time. When you see the model burning too much time on one planet scene, manually adjust its priorities to pull attention back to the core experience.

Step 3: Manual cleanup + persist the experience as a skill

After stopping the auto loop:

  1. Use another Claude session to fix rendering issues, clean up the code, and finish deployment
  2. Ask Opus 5 to consolidate the lessons into a skill and commit it to the GitHub repo

Per Anshu's open-source repo, during the review process Opus 5 built its own acceptance tooling: 17 mandatory validation checks built in, automatic captures of different scenes for comparison, frame-rate statistics, and exposure data analysis. This was a byproduct the model proactively produced under pressure from the Judge Agent—and you can reuse this acceptance script directly in your own projects.

Results

The finished product should meet:

  • Smooth playback in the browser (60 fps target)
  • All textures, animation, and sound AI-generated, no external assets
  • All 17 checks of the built-in acceptance tooling passing
  • Experience persisted as a reusable skill (GitHub repo)

Open longsilence.anshu.dev to experience the finished result of this process yourself.

The Portable Core Pattern: Making AI Its Own Client

The Challenge Loop's real value isn't game development—it's a portable multi-agent verification pattern:

Main Agent (splits tasks)
  ├── Builder Agent × N (works in parallel)
  └── Judge Agent (independent, blind-tested, no concessions)
        ↓
   Pass? No → rework
        Yes → next iteration

This structure fits any long-horizon task with an objective reference:

  • Web development: the Judge Agent blind-compares your finished build against competitor screenshots
  • Document writing: the Judge Agent compares against top industry articles item by item
  • Code refactoring: the Judge Agent accepts against test cases and performance benchmarks as hard metrics

The community has already proven its portability: Ryan Campbell iterated a go-kart game with the same prompt, and Yogi Suria went straight to building Claudepunk 2077. The same instructions also run on GPT-5.6 Sol, performing only slightly below Opus 5.

FAQ

  • What if costs explode: trial-run on a small goal first (2–4 hours); extend to 24 hours only after the Judge Agent's acceptance bar proves reasonable
  • What if the Judge Agent goes soft: explicitly require "blind side-by-side comparison" in the prompt and forbid the model from lowering its own bar; reference the 17-check script approach in Anshu's repo
  • The model grinds on one detail forever: step in manually and adjust priorities—exactly what Anshu himself did repeatedly over the 24 hours