Sakana Fugu: the Orchestrator That Doesn't Answer Itself, Just Directs Other Models to Do the Work

·Toolin Editorial Team

Sakana AI releases the Fugu family of orchestrator models, smartly dispatching GPT, Claude, and Gemini to finish tasks, with performance approaching Fable 5 and Mythos Preview.

Sakana Fugu: the Orchestrator That Doesn't Answer Itself, Just Directs Other Models to Do the Work

We're used to asking "which model is strongest." Sakana AI poses a new question: how do you make multiple cutting-edge models stronger together? On June 22, the Japanese AI unicorn co-founded by Llion Jones, fifth author of the Transformer paper, released the Sakana Fugu family of orchestrator models — it doesn't answer questions itself; it decides which model to hand the task to, then synthesizes the answer.

It suits developers who want "the best model per task," enterprises that want to avoid being locked to a single vendor, and researchers interested in multi-agent orchestration.

The Sakana Fugu series: many "little fish" converging into one "big pufferfish"

What Sakana Fugu Is

The official animation's metaphor for Fugu (pufferfish) is direct: many "little fish" converge into one "big pufferfish," that prized delicacy. Mapped onto models —

  • Don't train a stronger base model to solve problems; train a "commander-in-chief" that learns when to use which model
  • Fugu itself is a language model specialized in understanding "when to delegate a task, how agents communicate, and how to integrate the work into a reliable answer"
  • This route builds on the team's earlier research on learned model orchestration, including the ICLR 2026 papers Trinity and Conductor

Sakana AI argues in its blog that orchestration models will surpass traditional large models to become the new frontier. The reasoning: complex tasks demand expertise far beyond any single model's capability boundary, and getting the best performance out of models requires collective intelligence.

Core Mechanism: Four Command Actions

The technical report breaks Fugu's working mechanism into four steps:

1. Identify the problem type Judge whether the user's question is code, math, reasoning, information retrieval, scientific analysis, or a multimodal task. This step sets the starting point for the entire dispatch logic that follows.

2. Pick the right worker model Different models vary widely across tasks. One of Fugu's training goals is learning "which model to call for which problem." The report notes in particular that even within one task category (say, competitive programming), different models may respectively excel at direct implementation, devising a solution plan, or combining multiple algorithmic ideas — and Fugu must factor these fine differences into its decisions.

3. Design the agent workflow For complex problems, Fugu Ultra generates a complete agentic workflow, including task decomposition, subtask assignment, context-sharing strategy, and final answer synthesis — all done inside the model in natural language.

4. Optimize from feedback Fugu's training goes beyond supervised fine-tuning to include evolutionary algorithms and reinforcement learning, using real task outcomes to optimize orchestration strategy in reverse.

Two Versions: Daily Use vs. Hard Problems

VersionPositioningOrchestration styleBest for
FuguEveryday use, balancing performance and latencyLightweight selection mechanism, fast worker decisionsHigh-frequency, latency-sensitive use
Fugu-UltraQuality firstComplex orchestration, multi-agent collaboration + synthesisComplex code, mathematical reasoning, science questions, multi-step planning

What the two share is fully model-agnostic modularity: Fugu doesn't need access to the worker models' weights — they don't even need to be open source. New models can join the worker pool as soon as they ship, and users can customize the available model list by cost, privacy, compliance, and other needs.

Benchmarks: Beating Fable 5 and Mythos Preview on Three of Them

The technical report lists how the Fugu family performs across eight benchmarks covering four dimensions: coding, reasoning, science, and agent capability.

Fugu beats Mythos Preview and Fable 5 on three benchmarks

According to the report, purely through smart dispatch, the Fugu models beat Mythos Preview and Fable 5 on three benchmarks.

Cross-domain adaptability is also intuitive:

  • On Terminal Bench (terminal engineering tasks), the peak model both Fugu and Fugu Ultra called was GPT-5.5, the top performer in that benchmark
  • On GPQA Diamond (graduate-level science reasoning), Gemini-3.1-Pro was the leading model, and both Fugu models centered their dispatch on Gemini

In other words, Fugu doesn't try to replace GPT, Claude, or Gemini — it combines their capabilities.

A Few Interesting Experiments

The technical report's appendix has three experiments that show the orchestration ability vividly:

One-shot Rubik's Cube solver: the model must write, in one shot, a cube-solving program implemented with the Python standard library, tested on 300 scrambled cubes. Both Fugu and Fugu-Ultra solved every cube — Fugu-Ultra with a shorter average move count, Fugu running faster.

Blindfold chess: with no visible board, no legal-move list, and no FEN, the model continues a game from the move history alone — mainly a test of maintaining internal state over the long run. In a representative game, Fugu beat multiple baseline models and a handicapped Stockfish.

Online stock trading: the model sees only past and current anonymized market data, cannot peek at future prices, and must make buy/hold/sell decisions week by week. Fugu-Ultra posted a higher average return across five runs.

Netizens also threw Fugu-Ultra some classic traps that break many models — "how many r's in strawberry," "is 5.11 bigger than 5.1," the classic car-wash riddle — it got all three right.

Resources and Where to Try It

Use Cases and Caveats

Who it's for:

  • Developers who want to auto-pick the best model per task type and reduce single-vendor dependence
  • Enterprises that care about "AI sovereignty" and worry about export controls cutting off supply
  • Researchers in multi-agent collaboration and model orchestration

Caveats to watch:

  • Higher cost and latency: multi-model orchestration is naturally pricier and slower than a single model — most visible in Fugu-Ultra's deep-collaboration mode
  • Error attribution is complex: when the final answer is wrong, it's hard to tell whether routing, the worker model, or the synthesis step is at fault
  • The orchestrator itself carries bias risk: if it misjudges the task type or over-relies on one model, overall performance degrades

💡 Tip: The direction Fugu proposes expands "AI competition" from "single-model capability" to "system-organization capability" — whoever is better at dispatching models, using tools, designing workflows, and integrating feedback holds the greater power. The route has real potential, but real adoption still demands plenty of engineering validation. The test results in the technical report come from the vendor; actual capability will show in real developers' hands-on experience.

Related articles

TerminalWorld: The First Agent Benchmark Built on Real CLI Workflows
AI Tutorials

TerminalWorld: The First Agent Benchmark Built on Real CLI Workflows

A 1,530-task benchmark distilled from 80,000 human terminal recordings, spanning 18 workflow categories and 1,280 command tools — a cure for Agents that rack up leaderboard points yet fall apart in real terminal scenarios.

Toolin Editorial Team
Ruoyu Lanyue 01: The World's First AI Explosion-Proof Robot Fuels Real Cars by Itself
AI Products

Ruoyu Lanyue 01: The World's First AI Explosion-Proof Robot Fuels Real Cars by Itself

Driven by the Ruoyu Jiutian robot brain, the explosion-proof Lanyue 01 robot autonomously runs the full workflow at gas stations 24/7 — opening the fuel door, grabbing the nozzle, fueling, returning the nozzle — bringing embodied intelligence into high-risk environments.

Toolin Editorial Team
TRIAD: Teaching AI Agents Not Just to Refuse, but to Repair Dangerous Plans
AI Products

TRIAD: Teaching AI Agents Not Just to Refuse, but to Repair Dangerous Plans

The open-source Agent safety framework TRIAD replaces binary guardrails with three-way decisions (proceed/update/refuse), preserving the user's original task even under prompt injection attacks.

Toolin Editorial Team
Volcano Engine's Full Agent Infra Upgrade: The 1+N+X System, AgentKit, and ArkClaw Enterprise
AI Products

Volcano Engine's Full Agent Infra Upgrade: The 1+N+X System, AgentKit, and ArkClaw Enterprise

From the Seedance 2.0 moment to enterprise-grade Agent infrastructure, Volcano Engine uses the 1+N+X system, AgentKit, and ArkClaw Enterprise to move Agents from personal tools into organizational workflows for real.

Toolin Editorial Team
Baidu DuMate Hands-On Guide: A Homegrown Codex That Lets You Run Office Work by Voice
AI Tutorials

Baidu DuMate Hands-On Guide: A Homegrown Codex That Lets You Run Office Work by Voice

From installation to automation, a complete breakdown of Baidu DuMate's request -> authorize -> execute -> deliver pipeline: cross-app work, scheduled tasks, and the credits bill, all explained in one article.

Toolin Editorial Team
DeNovoSWE: The First Long-Horizon Doc2Repo Training Set That Teaches Code Agents to Build Repositories
AI Products

DeNovoSWE: The First Long-Horizon Doc2Repo Training Set That Teaches Code Agents to Build Repositories

Renmin University of China releases an open training set of 4818 real task instances targeting repository-level code generation, lifting Qwen3-30B's pass rate from 5.8% to 47.2%.

Toolin Editorial Team