OpenAI says Ultra Fast runs at eight times standard speed
At DevDay, a presenter said the service would reach API, ChatGPT and Codex; OpenAI's public materials describe earlier speed tiers with different multipliers.
By Ryan Merket · Published
Primary source: YouTube
Why it matters
OpenAI is extending its speed-tier pitch beyond the API, but developers cannot assess the tradeoff from an eight-times headline alone: Ultra Fast pricing, access conditions and measurement details were not specified, while OpenAI's published tiers cite different multipliers.

OpenAI said it is launching Ultra Fast across its API, ChatGPT and Codex, with a presenter at its September 29th DevDay keynote describing the tier as eight times the speed of Standard processing. The claim appeared in the keynote livestream; OpenAI's public documentation describes earlier speed tiers with different performance figures, and does not yet independently confirm the new rollout or multiplier.
The presenter contrasted Ultra Fast with Fast, which they described as twice as fast as Standard at twice the price. OpenAI's Fast mode documentation instead says Fast can deliver up to 2.5 times Standard speed for supported API workloads. It describes a paid service tier, configured per request or at project level, rather than a different model. OpenAI cautions that actual speed depends on conditions: requests share rate limits with Standard, and traffic that ramps too quickly may be served at Standard speed instead.
The distinction is important for developers evaluating latency and cost. A headline multiplier does not establish how quickly a particular application will respond. Model output speed is one component; time to first token, network overhead, tool calls and queueing also shape the user-facing delay. OpenAI's documentation says Fast carries a per-token premium. The keynote's two-times-price comparison concerned Fast; the presenter did not give Ultra Fast pricing, a latency guarantee, eligibility requirements or a measured workload for the eight-times figure.
There is another public benchmark to reconcile. In an August 13th announcement, OpenAI described an API-only limited preview of Ultrafast for GPT-5.6 Sol, saying it ran up to 14 times faster than Standard processing and could generate up to 750 output tokens per second. OpenAI attributed that service to its partnership with Cerebras and said access was limited to selected customers while capacity expanded. The August figure and Tuesday's eight-times statement refer to announcements with different dates and potentially different configurations; the keynote did not explain how they compare.
The distribution claim is the more concrete change in the September presentation: the speaker named API, ChatGPT and Codex, extending the announced reach beyond the earlier API preview. OpenAI's API model documentation treats processing speed as a service choice attached to a model, and its pricing page lists separate Fast rates for supported models. Those public pages provide a way to assess cost for documented tiers, but do not establish Ultra Fast's price or which ChatGPT plans and Codex users will receive it.
OpenAI has spent the year presenting low latency as infrastructure for interactive AI, rather than simply a faster model label. In February, the company introduced Codex-Spark, a coding model designed for real-time use and served on Cerebras hardware. OpenAI said Spark exceeded 1,000 tokens per second and paired the specialized inference path with changes to its API request pipeline. That launch focused on Codex and a restricted preview; it does not verify that the Ultra Fast service announced at DevDay uses the same hardware or reaches the same speed in ordinary workloads.
For API developers, the commercial test will be whether faster responses justify the service premium in applications where people are waiting on the model. For ChatGPT and Codex users, the announcement names the products but leaves practical access and pricing unspecified in the presentation captured here. Until OpenAI publishes the corresponding terms and measurement details, eight times is the presenter's comparison, not a performance guarantee for every request.