OpenAI has unveiled Ultrafast mode for GPT-5.6 Sol, its most capable model, running at up to 750 output tokens per second on Cerebras wafer-scale hardware—roughly 5× the ~150 tokens per second typical of production models and 10× faster than standard NVIDIA H100 deployments. The tier is available in limited preview today via the OpenAI API, with access expanding as capacity grows.
The speed breakthrough matters most for agentic workflows that chain dozens of model calls. A multi-step operation that would take 8 seconds per step on standard hardware now completes in ~2 seconds per step. For enterprises running thousands of agent runs daily, this compounds into workflows shifting from multi-minute completions to seconds—enabling real-time incident response, financial analysis, customer support, and commerce use cases where latency currently blocks adoption. OpenAI is already using Ultrafast internally for incident response and research workflows; early customers include Jane Street and companies across coding, commerce, and financial services.
Cerebras' partnership underpins the feat: OpenAI and Cerebras formalized a multi-year agreement in January 2026 to deploy 750 megawatts of dedicated low-latency inference capacity. The deployment is strategically timed: Cerebras is preparing for a 2026 IPO, and landing OpenAI as a reference customer strengthens its growth narrative ahead of roadshow discussions.