Anthropic releases Claude Opus 4.8, hailed as the 'most honest' model yet and featuring a new dynamic workflow tool.

Anthropic's Claude Opus 4.8 introduces a dynamic workflow tool and is noted for its improved honesty in handling errors. As the competition with OpenAI intensifies, this model is positioned as the most honest AI assistant yet.

Sources:
Yahoo FinanceAxios+5
Trending 22m ago
Tab background
Sources: AxiosThe Verge9to5Mac+1
Anthropic’s Claude Opus 4.8, unveiled on Thursday, amplifies the company’s commitment to transparency and effectiveness in AI. This upgraded model is said to have sharper judgment and is significantly less likely to make unsupported claims, marking a notable shift in AI behavior.

Current pricing remains unchanged at $5 per million input tokens, yet the enhancements are profound. According to early testers, Opus 4.8 is around 4x less likely to pass code flaws without notice compared to its predecessor, Opus 4.7.

“We’re making swift progress on developing these safeguards,” Anthropic reported, expressing hope for the rollout of Mythos-class models in the coming weeks. The new model encompasses a fast mode capable of operating at 2.5 times the speed of previous versions, with a cost reduction of three times compared to earlier models.

On numerous benchmarks, Opus 4.8 outperformed competitors, showcasing a coding score of 69.2% in SWE-Bench Pro, surpassing Opus 4.7’s 64.3%. Knowledge work score leaped from 1753 to 1890.

“Opus 4.8’s tendency to proactively flag issues with the inputs and outputs of an analysis, something other models routinely missed,” noted a testimonial from Bridgewater Associates. This new model ensures that AI systems prioritize transparency, addressing a longstanding issue where AI systems could confidently make claims without substantial evidence.

The upgrade cycle has accelerated significantly; just 41 days after the launch of Opus 4.7, Anthropic is positioning itself as a leader in producing reliable AI tools with enhanced capabilities for users.
Sources: Yahoo FinanceThe Verge
Anthropic has launched Claude Opus 4.8, the latest version of its AI model, boasting improved honesty and coding capabilities. Priced similarly to its predecessor at $5 per million input tokens, the model is designed to better flag uncertainties and avoid unsupported claims, enhancing user trust in AI tasks.
Section 1 background
The Headline

Anthropic releases Claude Opus 4.8

Key Facts
  • Anthropic released Opus 4.8, the newest version of its most advanced publicly available model, on May 28, 2026.TechCrunch
  • Opus 4.8 was launched just 41 days after the release of Opus 4.7, marking a much faster upgrade cycle than normal for Anthropic.TechCrunch
  • Claude Opus 4.8 brings improvements in coding and honesty, while Anthropic says Mythos-class models could reach the broader public within weeks.Yahoo Financeqz.com
  • Benchmarks in coding, agentic tasks, reasoning, and practical knowledge work all showed gains over Opus 4.7, Anthropic said.qz.com
  • Opus 4.8 is roughly four times less likely than Opus 4.7 to allow flaws in code it has written to pass without comment.qz.com
  • Anthropic launched a new feature alongside Opus 4.8, designed to help manage complex tasks across hundreds of parallel subagents.TechCrunch
  • Users will have control over how much effort Claude puts into specific tasks with a dropdown menu selecting effort levels.inc.com
  • The company hinted that the Mythos preview period might soon end, once necessary safeguards are in place.
Section 2 background
Background Context

Background on Claude Opus 4.8

Key Facts
  • Anthropic describes Claude Opus 4.8 as having “sharper judgement, more honesty about its progress, and the ability to work independently for longer than its predecessors.”9to5Mac
  • Opus 4.8 achieved a record score of 69.2 percent in SWE-Bench Pro for coding abilities, surpassing Claude Opus 4.7's 64.3 percent and OpenAI's GPT-5.5's 58.6 percent.inc.com
  • Opus 4.8 scored 1890 in an OpenAI-created benchmark for economically-viable work, up from Opus 4.7's 1753 and GPT 5-5's 1769.inc.comAxios
  • Agentic coding score increases from 64.3% to 69.2%.9to5Mac
  • Multidisciplinary reasoning with tools jumps from 54.7% to 57.9%.9to5Mac
  • Agentic computer use moves from 82.8% to 83.4%.9to5Mac
  • Knowledge work score increases from 1753 to 1890.9to5Mac
  • Agentic financial analysis improves from 51.5% to 53.9%.9to5Mac
  • Anthropic offered a broader critique of AI behavior, quoting its own documentation: "sometimes jump to conclusions, confidently claiming to have made progress in their work despite the evidence being thin."qz.com
  • Opus 4.8 reaches new highs on measures of what Anthropic called prosocial traits, including supporting user autonomy and acting in users' best interests.qz.com
Article not found
CuriousCats.ai

Article

Source Citations