OpenAI's Ultrafast mode runs GPT-5.6 Sol up to 14x faster using Cerebras wafer-scale hardware, now in limited API preview.
TECH

OpenAI's Ultrafast mode runs GPT-5.6 Sol up to 14x faster using Cerebras wafer-scale hardware, now in limited API preview.

20+
Signals

Strategic Overview

  • 01.
    OpenAI is previewing Ultrafast, a new service tier for GPT-5.6 Sol that runs inference up to 14x faster than standard processing, launching first in the OpenAI API for a limited group of customers.
  • 02.
    Ultrafast generates up to 750 output tokens per second, powered by Cerebras' Wafer-Scale Engine hardware, which packs 44 GB of SRAM directly on each wafer-sized chip.
  • 03.
    The joint announcement came August 13, 2026; availability is currently a limited preview with wider access planned as capacity grows, and businesses can join a waitlist by sharing workload, latency, and usage details.
  • 04.
    On the GDP-Val benchmark of economically valuable knowledge-work tasks, such as legal briefs and financial models, Ultrafast delivered a 5.6x end-to-end speedup with no loss in quality.

Deep Analysis

The Wafer-Scale Trick That Makes 750 Tokens a Second Possible

Ultrafast's headline number - up to 14x faster inference, topping out around 750 output tokens per second [1]- isn't achieved by shrinking the model. It's the same GPT-5.6 Sol, run on hardware that sidesteps the bottleneck that normally throttles large language models: memory bandwidth. On a standard GPU cluster, weights have to be shuttled between off-chip memory and processing cores for every token generated. Cerebras' Wafer-Scale Engine instead keeps 44 GB of SRAM directly on each wafer-sized chip [1], so the model's weights sit next to the compute rather than queuing behind a bus. OpenAI frames the result as a tradeoff-free kind of scaling: 'Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.' [2]

The benchmark numbers back up the framing. On GDP-Val, a suite of economically valuable knowledge-work tasks like legal briefs and financial models, Ultrafast delivered a 5.6x end-to-end speedup with no measured loss in quality [1]. On Humanity's Last Exam, a 2,500-question benchmark, Ultrafast finished the full set in 11 hours 11 minutes versus 78 hours 27 minutes for Anthropic's Claude Fable 5 at comparable accuracy - roughly a 7x wall-clock difference [1]. That gap matters more than the raw tokens-per-second figure: it's the difference between a benchmark run that finishes overnight and one that finishes before lunch.

Speed Becomes a New Axis in the Anthropic Rivalry

Cerebras and OpenAI didn't just publish a speed number in isolation - they published it against Anthropic by name. Citing Artificial Analysis output-speed data, the companies say Ultrafast runs roughly 5x faster than Claude Opus 4.8 in Fast mode and 11x faster than Claude Fable 5 [1]. Cerebras CEO Andrew Feldman put it plainly: 'GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive.' [1]Coverage from TechCrunch [3]and 9to5Mac [4]both led with the 14x figure, a sign that inference speed, not just benchmark accuracy, is becoming a marketable differentiator between frontier labs.

This didn't happen overnight. OpenAI and Cerebras signed a multi-year deal back in January 2026 to deploy up to 750 megawatts of wafer-scale systems into OpenAI's inference stack, phased through 2028 [5], [6]. Ultrafast is the first visible product of that infrastructure bet, arriving seven months later alongside a formal press release [7]. The timing suggests OpenAI viewed raw generation speed as a strategic gap worth a dedicated hardware partnership, rather than something to solve purely in software.

The Real-World Case: From Hour-Long Waits to Before You Look Away

OpenAI's pitch for Ultrafast centers on workflows that break down when a model takes too long to answer: voice applications, customer support, e-commerce, developer agents, financial research, incident response, and cybersecurity threat detection [2]. Internally, OpenAI says incident-response investigations that used to take one to two hours now resolve in ten to fifteen minutes on Ultrafast. One researcher described responses returning before he context-switches away from a task [1]- a small but telling detail, since context-switching cost, not raw compute, is often the real productivity tax of slow AI tools.

Reception outside the official channels leans toward validating that framing, with a caveat. Early testers describe GPT-5.6 Sol on Ultrafast as a fast, dependable workhorse rather than a leap in raw intelligence, and note it is the first model in the 5.6 line trustworthy enough to run autonomous sub-agent tasks unsupervised. That distinction matters for where the speedup pays off: as more workflows involve parallel sub-agents firing off many small requests rather than one long conversation, the bottleneck shifts from how smart the model is to how fast it can respond - exactly the gap Ultrafast is built to close.

Preview Now, Price Tag Later

Everything about Ultrafast so far is described as a preview: available first through the OpenAI API to a select group of customers, with no public pricing and wider access planned as capacity grows [2]. Businesses that want in have to join a waitlist and describe their workload, latency requirements, and expected usage [4]. That gating matters because speed like this is presumably expensive to serve - wafer-scale chips are a scarcer, costlier resource than commodity GPUs - and neither company has said what a 14x speedup will cost end users.

That silence is where public reaction turns skeptical. Absent official pricing, discussion elsewhere has ranged from genuine enthusiasm about the underlying hardware achievement to pointed cost anxiety and even suspicion that standard-mode responses might be paced deliberately slower to make a premium fast tier look more valuable - a claim neither OpenAI nor Cerebras has addressed. Others argue the more interesting bottleneck isn't model speed at all: as parallel sub-agent workflows spread, local infrastructure like compilers and CPU throughput could become the next constraint regardless of how fast the model itself responds. Both critiques point to the same open question the preview leaves unanswered - whether Ultrafast is a genuine step change in how AI gets used, or a capacity-constrained showcase whose real test only comes once pricing and availability go broad.

Historical Context

2026-01-14
OpenAI and Cerebras signed a multi-year deal to deploy up to 750 megawatts of Cerebras wafer-scale systems into OpenAI's inference stack, phased through 2028, laying the infrastructure groundwork for Ultrafast.
2026-06
OpenAI introduced the GPT-5.6 model family (Sol, Terra, Luna), with Sol positioned as the most capable variant and Luna as the speed-focused one.
2026-08-13
OpenAI and Cerebras jointly announced the Ultrafast preview for GPT-5.6 Sol, publishing coordinated blog posts and a press release.

Power Map

Key Players
Subject

OpenAI's Ultrafast mode runs GPT-5.6 Sol up to 14x faster using Cerebras wafer-scale hardware, now in limited API preview.

OP

OpenAI

Developer of GPT-5.6 Sol and the Ultrafast mode; controls API access and rollout pace, positioning Ultrafast for latency-sensitive workflows like voice, customer support, developer agents, and financial research.

CE

Cerebras Systems

Hardware and infrastructure partner powering Ultrafast via its Wafer-Scale Engine, part of a broader multi-year deal to deploy up to 750 megawatts of wafer-scale systems into OpenAI's inference stack through 2028.

AN

Anthropic

Competitor whose Claude Opus 4.8 and Claude Fable 5 models are used as the explicit speed-benchmark comparison for Ultrafast's marketing claims.

AR

Artificial Analysis

Independent benchmarking firm whose output-speed comparisons are cited as the basis for the 5x and 11x speed claims against Claude models.

Fact Check

7 cited
  1. [1] Accelerating GPT-5.6 Sol Ultrafast with OpenAI
  2. [2] Previewing Ultrafast
  3. [3] OpenAI introduces Ultrafast, a new mode that makes GPT-5.6 Sol work at 14x the speed
  4. [4] OpenAI previews Ultrafast GPT-5.6 Sol, running up to 14 times faster
  5. [5] Cerebras Lands Major OpenAI Deal to Scale AI Inference
  6. [6] OpenAI Partners with Cerebras to Deploy 750MW Wafer-Scale Systems for High-Speed Inference
  7. [7] Cerebras Powers Ultrafast Mode for OpenAI's GPT-5.6 Sol

Source Articles

Top 1

THE SIGNAL.

Analysts

Frames Ultrafast as evidence that raw model intelligence and inference speed are no longer a tradeoff: 'GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive.'

Andrew Feldman
CEO and co-founder, Cerebras

Positions Ultrafast as matching AI response speed to the pace of human thought and coding workflows: 'With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code.'

Rohan Varma
Product, OpenAI

Describes a first-hand workflow improvement where Ultrafast responses return before he context-switches away from a task: 'Whereas formerly I might have to wait a couple minutes, it now finishes before I context-switch.'

Jeffrey Wang
Researcher, OpenAI
The Crowd

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows.

@@OpenAI11329

Previewing Ultrafast mode for @OpenAI's GPT 5.6 Sol, powered by Cerebras. GPT-5.6 Sol Ultrafast generates responses at up to 750 tokens per second. That's the full, GPT-5.6-Sol model -- up to 14x faster than the same model on Standard processing. It speedran Humanity's Last

@@cerebras3226

GPT-5.6-Sol Ultrafast mode (750 TPS) is here powered by Cerebras! Here's How kernels work on Cerebras Chips cerebras built a chip with 900,000 cores on a single silicon wafer. CSL (Cerebras Software Language) is a Zig-inspired DSL that gives you direct control over the

@@wafer_ai509

Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed

@u/YeXiu223217
Broadcast
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

I Tested GPT-5.6 Sol for a Month

I Tested GPT-5.6 Sol for a Month

Finally GPT 5.6 Sol Ultra Inside Codex (750 Token/Sec) | GPT 5.6 sol updates | GPT 5.6 + Codex

Finally GPT 5.6 Sol Ultra Inside Codex (750 Token/Sec) | GPT 5.6 sol updates | GPT 5.6 + Codex

OpenAI's Ultrafast mode runs GPT-5.6 Sol up to 14x faster using Cerebras wafer-scale hardware, now in limited API preview. — AI News | Agentic Brew