- Sakana has launched an AI system called Fugu that reportedly beats Anthropic's Claude 5 on certain benchmarks.
- Fugu Ultra matched or exceeded Anthropic's Fable 5 and Mythos on key engineering and science benchmarks.
- Fugu outperformed Claude Fable 5 on coding and Mythos Preview on graduate-level science multiple-choice tests.
- Benchmark charts shared by Sakana show, Fugu exceeds the performance of Anthropic's Claude Fable 5 on LiveCodeBench, an open source benchmark testing coding performance on regularly refreshed, software problem-solving tasks (Fugu Ultra: 93.2, Fugu: 92.9, Fable: 89.8).
- Fugu beats the prior Claude Mythos Preview model on GPQA-D (Diamond), a test of 198 graduate-level multiple-choice questions in biology, physics, and chemistry (Fugu Ultra: 95.5, Fugu: 95.5, Mythos Preview: 94.6).
- Fugu was launched in two versions: Fugu for coding, chat, and other everyday tasks, and Fugu Ultra for more complex work such as AI research, paper reproduction, cybersecurity analysis, and patent investigations.
Sakana's Fugu AI system has made waves in the AI community by reportedly outperforming Anthropic's Claude 5 on several key benchmarks. The company launched two versions of Fugu: the standard model for everyday tasks and the advanced Fugu Ultra for complex applications.123456
Benchmark data indicates that Fugu Ultra matched or exceeded Claude's performance on critical engineering and science tests. For instance, on the LiveCodeBench, which evaluates coding performance, Fugu Ultra scored 93.2, while Fugu scored 92.9, compared to Claude Fable's 89.8.
In graduate-level science assessments, Fugu also outperformed Claude Mythos Preview on the GPQA-D (Diamond) test, achieving a score of 95.5 in biology, physics, and chemistry, compared to Mythos Preview's 94.6.
These advancements come as Anthropic's Fable 5 and Mythos 5 models faced a rollback shortly after their launch due to national security concerns raised by the US government. Despite this, Vals AI ranks Fable 5 as the most capable publicly available AI model based on its benchmark tests. Sakana's Fugu, however, is positioning itself as a formidable competitor in the rapidly evolving AI landscape.
“Sakana has launched its Fugu AI system, which reportedly outperforms Anthropic's Claude 5 on several benchmarks. The Fugu Ultra version is designed for more complex tasks, showcasing significant advancements in AI capabilities.”