- Claude Fable 5 has surpassed OpenAI GPT 5.5, Google Gemini 3.5 Pro, and Claude Mythos Preview across multiple benchmarks, demonstrating superior performance.
- On SWE-Bench Pro, which evaluates real-world software engineering performance, Fable 5 achieved 80.3%, ahead of Claude Mythos Preview’s 77.8%, GPT 5.5’s 58.6%, and Gemini 3.5 Pro’s 54.2%.
- On GDPval-AA, Anthropic’s benchmark for complex analytical tasks, Fable 5 scored 1932, placing it ahead of Mythos Preview at 1869, GPT 5.5 at 1769, and Gemini 3.5 Pro at 1314.
- On GDPpdf, evaluating document and visual reasoning capabilities, Fable 5 scored 29.8%, compared with 24.9% for GPT 5.5 and 16.7% for Gemini 3.5 Pro.
- Fable 5 also scored 78.0%, compared with 69.0% for Mythos Preview and 34.0% for GPT 5.5.
Anthropic's Claude Fable 5 has significantly surpassed competitors in a series of benchmarks. On SWE-Bench Pro, a vital measure of software engineering capabilities, Fable 5 achieved 80.3%, outpacing Claude Mythos Preview at 77.8%, OpenAI's GPT 5.5 at 58.6%, and Google Gemini 3.5 Pro at 54.2%.12
In complex analytical tasks, Fable 5 scored 1932 on GDPval-AA, leading over Mythos Preview's 1869, GPT 5.5's 1769, and Gemini 3.5 Pro's 1314.3
When tested on GDPpdf, assessing document and visual reasoning, Fable 5 achieved 29.8%, compared to 24.9% for GPT 5.5 and 16.7% for Gemini 3.5 Pro. In another assessment, Fable 5 scored 78.0%, versus 69.0% for Mythos Preview and 34.0% for GPT 5.5.45
These results illustrate Anthropic's advancements, reinforcing Claude Fable 5's position at the forefront of AI technology across multiple domains, including coding, reasoning, vision, cybersecurity, biology, and automation tasks.
“Claude Fable 5 has demonstrated strong performance against its competitors, scoring 80.3% on SWE-Bench Pro and outpacing others in complex analytical tasks. Its benchmark results across various metrics highlight its superiority in the AI landscape.”
