OpenAI introduced Ultrafast, a preview service tier that runs GPT-5.6 Sol up to 14 times faster than standard processing and generates up to 750 output tokens per second.
Key Points:
- Ultrafast runs GPT-5.6 Sol up to 14 times faster than standard processing.
- Cerebras powers the mode, which can deliver up to 750 output tokens per second.
- The service is in limited preview, with broader access planned as capacity grows.
OpenAI Ultrafast Preview
OpenAI announced Ultrafast on Aug. 13, launching the speed-focused tier first through its API for a limited group of customers. The company said access will expand as capacity grows.
The service is powered by Cerebras and is designed to preserve GPT-5.6 Sol’s capabilities while cutting the time required to produce responses. Output tokens are pieces of generated text. OpenAI said real-time performance had typically required users to choose a smaller or more specialized model, while Ultrafast targets “more useful work per second.”
OpenAI listed incident response, financial research, customer support, commerce and live research among the workflows it expects to benefit most. Engineers can use the mode to review logs, traces and recent code changes while an outage is unfolding, while support systems can handle multi-step requests during live conversations.
Also Read: Anthropic Pursues A $6B Decart Deal While Preparing To Go Public
GPT-5.6 Sol Impact
Early customers said the speed changes how the model can fit into interactive work, with John Crepezzi of Jane Street calling the increase “impressive” and saying it makes focused work alongside models more practical.
Courtland Lykins, product lead for voice AI at Podium, said Ultrafast has been valuable for the company’s voice stack because “the speed completely changes the call experience” on complex work. OpenAI also said financial researchers can analyze changing market signals and transactions without waiting for slower batch-style processing. That use case depends on rapidly changing inputs.
The preview remains constrained by infrastructure rather than a broad public rollout. OpenAI said it is using feedback from the initial companies to determine where the speed produces the most value before expanding availability.
The Cerebras connection predates Ultrafast. OpenAI announced a partnership with the chipmaker in January to add 750 megawatts of low-latency AI compute to its platform. In February, GPT-5.3-Codex-Spark became the first model tied to that partnership, using Cerebras hardware to exceed 1,000 tokens per second in real-time coding.
Read Next: Virtuals Protocol Trading Volume Explodes 180%, Bulls Eye $0.68





