OpenAI previews GPT-5.6 Sol at up to 14x standard speed

Cerebras hardware drives the API tier to 750 output tokens per second, with access restricted to selected customers while capacity grows.

By · Published · Updated

Primary source: OpenAI on X

Why it matters

OpenAI is making inference speed a distinct API product for voice, coding and operational systems where latency limits usage. Its reach now depends on how quickly Cerebras capacity enters production.

OpenAI previews GPT-5.6 Sol at up to 14x standard speed

OpenAI (@OpenAI) opened a limited preview of Ultrafast on August 13th, a new API service tier that OpenAI says can run GPT-5.6 Sol up to 14 times faster than its Standard processing tier.

https://x.com/OpenAI/status/2087947721936359705

poster=/api/storage/public-objects/tweet-videos/openai-gpt-5-6-sol-ultrafast-cerebras-preview-poster-38177545.jpg|Video from @OpenAI on X

The Ultrafast announcement puts OpenAI's flagship model on Cerebras inference hardware, generating as many as 750 output tokens per second. OpenAI is starting with selected customers and plans to widen availability as compute capacity grows.

That rollout turns inference speed into a separate product tier rather than reserving the fastest responses for smaller models. OpenAI is positioning Ultrafast for applications where users or automated systems are waiting on an answer: incident response, voice agents, financial research, customer support, commerce and interactive scientific work.

OpenAI's 750-token figure measures output generation and does not by itself establish the time required to produce the first token or complete a workflow involving tools and external systems. The 14x comparison is also OpenAI's measurement against its own Standard processing tier. Independent production data will depend on prompts, reasoning settings, output length and the surrounding application.

Cerebras moves into OpenAI's inference stack

Ultrafast is the first named API service to emerge from a much larger infrastructure agreement. On January 14th, OpenAI said it had partnered with Cerebras to add 750 megawatts of low-latency computing capacity to its platform in stages through 2028.

Cerebras builds systems around wafer-scale processors that place compute, memory and bandwidth on a single chip. OpenAI said in January that this architecture would be assigned to workloads where response speed matters, giving OpenAI another inference option alongside the conventional accelerator clusters used across the AI industry.

The August preview begins putting that capacity in front of customers. It also gives Cerebras, led by co-founder and CEO Andrew Feldman, a production showcase for wafer-scale inference on OpenAI's flagship model rather than a smaller or specialized model.

OpenAI originally disclosed the planned Cerebras deployment in its June 26th preview of GPT-5.6 Sol, when it said the 750-token-per-second service would launch in July. GPT-5.6 reached general availability on July 9th, while the Cerebras-backed tier entered limited preview on August 13th.

Cerebras stock falls despite the OpenAI preview

Cerebras shares traded lower on August 13th despite OpenAI putting the chipmaker's hardware behind the new Ultrafast tier. The decline does not establish how investors view the broader partnership, particularly because OpenAI had previously disclosed both the infrastructure agreement and the planned 750-token-per-second service.

The latest announcement marks the service's limited customer preview rather than a new contract between the companies. Its commercial significance will depend on wider availability, customer adoption and Cerebras' ability to bring additional capacity into OpenAI's inference stack.

OpenAI starts with business workflows

OpenAI named Jane Street, Podium, Basis and Rogo among the early Ultrafast testers. Those customers cover coding assistants, voice systems and financial research, areas where delays can directly limit how a product is used.

Podium product lead Courtland Lykins said Ultrafast had been valuable in its voice stack because "the speed completely changes the call experience" for complex work. Rogo, which builds software for financial research, told OpenAI that the tier made complex analysis feel closer to a real-time interaction.

OpenAI is also testing Ultrafast internally for incident response and research. Engineers are using it to examine logs, traces and discussions during outages, while researchers are trying to compress experiment cycles that previously ran overnight into repeated iterations during a working day, according to OpenAI.

These cases explain why OpenAI is releasing the tier through the API first. Developers can redesign an application around lower latency, while a faster chat response mainly changes how an existing interface feels. Voice agents can sustain a conversation with fewer pauses. Coding tools can return larger edits while a developer remains focused on the task. Incident-response systems can process changing evidence before an outage moves into its next phase.

The underlying model remains GPT-5.6 Sol, which supports a 1.05 million-token context window and up to 128,000 output tokens. OpenAI currently lists standard Sol API pricing at $5 per 1 million input tokens and $30 per 1 million output tokens.

Ultrafast's restricted release makes capacity the immediate constraint. OpenAI said the initial customers will help determine where the additional speed produces measurable value and how future products should use it. Wider distribution will depend on Cerebras capacity entering OpenAI's inference stack, making the rollout an early test of whether specialized hardware can support frontier models at commercial scale.

Disclosure: The author owns Cerebras stock.

Reader comments

Conversation for this story loads after sign-in.