- OpenAI and Cerebras have announced an early look at Ultrafast, a new service tier for GPT-5.6 Sol that runs up to 14x faster than standard processing, launching first in the OpenAI API.
- Ultrafast is powered by Cerebras and generates up to 750 output tokens per second, enhancing the performance of OpenAI's most intelligent model.
- Ultrafast is currently available in a limited preview to a select group of customers, with plans for broader access in the future.
- Benchmarks indicate that GPT-5.6 Sol on Ultrafast runs 11x faster than Fable 5 and 5x faster than Opus 4.8 on Fast mode.
- In evaluations, Ultrafast answered all 2,500 HLE questions in 11 hours and 11 minutes, significantly faster than Claude Fable 5, which took 78 hours.
OpenAI and Cerebras have launched Ultrafast, a new service tier for GPT-5.6 Sol, which operates up to 14 times faster than standard processing. This service, available in a limited preview, is designed to enhance AI responsiveness in various applications, generating up to 750 output tokens per second.123456
During the preview, select customers are testing Ultrafast to identify its most impactful applications. “The increase in speed brought by Cerebras is impressive,” said John Crepezzi from Jane Street, highlighting the practical benefits for developers.
Ultrafast is particularly beneficial in high-stakes environments, such as incident response, where engineers need to act quickly. “With Ultrafast, we see this loop tightening to support multiple iterations during the workday instead,” noted a user.

Benchmarking results show that GPT-5.6 Sol on Ultrafast mode answered 2,500 high-level questions in just 11 hours and 11 minutes, compared to Claude Fable 5, which took over 78 hours. This performance demonstrates Ultrafast's capability to process complex tasks significantly faster, achieving comparable accuracy nearly 7 times faster.
Cerebras' innovative Wafer-Scale Engine architecture powers this service, allowing for efficient data processing without quality compromise. “Ultrafast allows us to create synchronous experiences for users that were previously limited by intelligence,” said Mitch Troyanovsky, co-founder of Basis.
As Ultrafast expands, it promises to transform workflows across industries, enhancing productivity and decision-making speed.
“Ultrafast delivers up to 750 output tokens per second, and in Cerebras benchmarks it answered all 2,500 HLE questions in 11 hours and 11 minutes, nearly 7x faster than Claude Fable 5. The service is available in limited preview to select customers, with access expanding over time.”









