Sakana AI released Fugu Ultra v2 and Fugu Max, models that orchestrate other models

Tokyo-based Sakana AI released Fugu Ultra v2.0 and Fugu Max v1.0 on September 11, 2026, under the heading “Orchestrating the Pareto Frontier.” Neither model is a conventional foundation model trained end to end on a single objective. Fugu is instead a learned orchestrator: a system trained to take one incoming request, decide which models in a fixed pool of open-weight and specialized models should handle which part of it, and stitch the results back into a single answer, recursively calling instances of itself where useful. Sakana grounds the approach in two of its own ICLR 2026 papers - TRINITY, which assigns Thinker, Worker, and Verifier roles across a task, and the Conductor, trained by reinforcement learning to perform natural-language orchestration.

The performance claims are pointed directly at the frontier labs’ flagship models without including them in the pool being orchestrated. Sakana reports Fugu Ultra v2 taking the best or joint-best score on five of eight benchmarks - GDP.pdf, Chartography, DeepSWE, Toolathon, and Sakana’s internal SWEFish - and a top-two placing on seven of eight. The pointed detail is in the footnote: Sakana states that Fugu Ultra’s training cutoff is 2026-08-28 and that Claude Fable 5, Fable 5.1 and GPT-6 Astra are NOT in Fugu Ultra’s model pool, so whatever it beats, it beats without routing to it. One worked example makes the claim concrete: on a 50-week simulated stock-trading pipeline starting from 10,000 dollars, Sakana reports Fugu Ultra ending at 11,943.22 dollars, a 19.43 percent mean return, against under 15 percent for the frontier models it compares against. A cost-optimized sibling, Fugu Max, integrates NVIDIA’s Nemotron family into its routing pool through a collaboration with NVIDIA and takes the best overall score on six benchmarks including Terminal Bench 2.1, GPQAD, AA-LCR and AutomationBench.

Pricing is per-token on top of an orchestration layer rather than a single model’s inference cost: Fugu Ultra runs $5 per million input tokens and $30 per million output tokens ($10/$45 above 272K context), while Fugu Max runs a flat $2 input / $6 output per million tokens regardless of context length - a price Sakana positions as 40 to 60 percent below Sonnet 5, GPT 5.6 Terra and Kimi K3 on output. Both ship as a hosted, OpenAI-compatible API only (no weights release) through Sakana’s own console plus OpenRouter and Vercel, and neither is available in the EU or EEA. A third variant, Fugu Cyber v1.0, is aimed at security work and Sakana reports an 86.9 percent success rate on CyberGym and 72.1 percent on CTI-REALM.

The significance is architectural rather than a capability leapfrog: Sakana is betting that composing a fixed pool of existing open and specialized models with a trained router beats training one larger monolithic model, and that the bet can win on both benchmark score and unit cost simultaneously. Buyers should treat the cross-lab benchmark comparisons as Sakana’s own numbers on Sakana’s own chosen benchmarks - Chartography and DeepSWE are not the full picture of frontier-model capability - but the release is a real signal that orchestration-over-models is now a competitive product category, not just a research idea.

Sources

Last verified September 14, 2026