John CrepezziMitch TroyanovskyJeffrey WangCourtland LykinsRohan VarmaAlex WangRogoCerebrasCerebras SystemsPodiumOpenAIBasisJane Street

OpenAI and Cerebras preview Ultrafast, a new service tier that runs GPT-5.6 Sol up to 14x faster, launching first in the OpenAI API

OpenAI and Cerebras have unveiled Ultrafast, a new service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing. This innovation, available first in the OpenAI API, aims to enhance responsiveness in AI applications, generating up to 750 output tokens per second.

OpenAI+1 source13 August 2026 · 19:29 UTC
CuriousCats Full Story

OpenAI and Cerebras have launched Ultrafast, a new service tier for GPT-5.6 Sol, which operates up to 14 times faster than standard processing. This service, available in a limited preview, is designed to enhance AI responsiveness in various applications, generating up to 750 output tokens per second.123456

During the preview, select customers are testing Ultrafast to identify its most impactful applications. “The increase in speed brought by Cerebras is impressive,” said John Crepezzi from Jane Street, highlighting the practical benefits for developers.

Ultrafast is particularly beneficial in high-stakes environments, such as incident response, where engineers need to act quickly. “With Ultrafast, we see this loop tightening to support multiple iterations during the workday instead,” noted a user.

Benchmarking results show that GPT-5.6 Sol on Ultrafast mode answered 2,500 high-level questions in just 11 hours and 11 minutes, compared to Claude Fable 5, which took over 78 hours. This performance demonstrates Ultrafast's capability to process complex tasks significantly faster, achieving comparable accuracy nearly 7 times faster.

Cerebras' innovative Wafer-Scale Engine architecture powers this service, allowing for efficient data processing without quality compromise. “Ultrafast allows us to create synchronous experiences for users that were previously limited by intelligence,” said Mitch Troyanovsky, co-founder of Basis.

As Ultrafast expands, it promises to transform workflows across industries, enhancing productivity and decision-making speed.

Key Insight
“Ultrafast delivers up to 750 output tokens per second, and in Cerebras benchmarks it answered all 2,500 HLE questions in 11 hours and 11 minutes, nearly 7x faster than Claude Fable 5. The service is available in limited preview to select customers, with access expanding over time.”
Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed
CuriousCats Shorts-list
Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed
CuriousCats studied:
1
OpenAI
“Today, we’re sharing an early look at Ultrafast, a new service tier that runs GPT‑5.6 Sol up to 14× faster than Standard processing, launching first in the OpenAI API.”
OpenAI →
2
Cerebras
“Today, Cerebras and OpenAI are sharing an early look at , a new service tier launching first in the OpenAI API and powered by Cerebras.”
Cerebras →
Ask CuriousCats
What is OpenAI's new Ultrafast service tier?
How does Ultrafast enhance GPT-5.6 Sol performance?
Who are the select customers for Ultrafast preview?
Are other AI models achieving similar speeds?
How does Ultrafast compare to Claude Fable 5?
Get your CIA-level briefing,
in real time.
CuriousCats monitors the internet every minute for you and brings you the most personalized brief of videos, social media posts, news and more.
Download the App
One story brought you here.
CuriousCats brings you everything else worth knowing.
Get CuriousCats