OpenAI says Responses API usage grew 100x in a year

At DevDay, a presenter paired the growth claim with reliability above 99%, without defining either the usage measure or reliability metric.

By · Published

Primary source: YouTube

Why it matters

Responses API growth could reflect expanding use of agent workflows, which make multiple calls for one task. Without a baseline or a defined reliability measure, OpenAI's 100x and above-99% claims do not show how much customer work the API handles or how consistently it completes it.

People of different ages collaborate around a laptop and shared notes, illustrating OpenAI’s mission for advanced intelligence to benefit humanity.

OpenAI said at its September 29th DevDay keynote that usage of its Responses API had grown 100-fold over the previous year, while reliability remained above 99%. The presenter gave no baseline for the growth figure and did not define what counted as reliable. The keynote video contains the claim, but the number is not accompanied by a chart or methodology in the available presentation material.

That leaves the headline growth rate hard to translate into customer adoption or business performance. OpenAI did not say whether the comparison measured API requests, completed responses, tokens, active developers, revenue, or another usage measure. It also did not specify whether the figure covered all Responses API customers and workloads, or a narrower segment. A 100-fold increase from a small starting point can describe a very different scale from the same multiple applied to an established service.

OpenAI's own March 11th retrospective said thousands of developers were using the API and described applications in customer support, legal services, life sciences and travel. It did not publish a request count or independently substantiate the 100-fold figure. The retrospective instead illustrates the range of software being built on the API, including multi-agent systems and workflows that use tools. OpenAI's account of the API's first year offers adoption examples, not a denominator for the DevDay claim.

The distinction between an API call and a customer's task is especially important here. Responses is built to support agent-like applications, with features including hosted tools, multi-turn state and multimodal inputs, according to OpenAI's API documentation. A task can involve repeated model requests as an agent chooses an action, runs a tool and sends the result back for another response. OpenAI's engineering team said a Codex coding workflow can involve dozens of back-and-forth Responses API requests for one task. A rising response count could therefore reflect more agent use, more steps per task, or both; the keynote did not separate those effects.

A customer task can involve repeated Responses API requests as an agent chooses an action, runs a tool, and sends the result back.
OpenAI says a Codex coding workflow can involve dozens of Responses API requests for one task — AI explanatory diagram, not documentary evidence. RuntimeWire · AI-generated diagram.

The reliability figure has a similar scope problem. OpenAI did not say whether its "more than 99 percent" referred to uptime, successful HTTP requests, completed responses, or a measure that also accounted for latency and partial failures. Its status documentation says availability metrics are aggregated across tiers, models and error types, and that an individual customer's availability can differ. That is a broad status measure, not necessarily a measure of whether a particular agent finished its work successfully.

A table distinguishes OpenAI’s 100-fold Responses API usage-growth claim, its undefined more-than-99-percent reliability claim, and the 99.9-percent Scale Tier uptime SLA.
The keynote figures have different scopes: OpenAI did not define its reliability measure, while the 99.9% SLA applies to Scale Tier traffic — AI explanatory infographic, not documentary evidence. RuntimeWire · AI-generated infographic.

OpenAI's incident records show why that definition matters. On June 2nd and 3rd, 2026, some Pay-As-You-Go and Priority Processing Responses API requests experienced elevated latency after a configuration rollout affected shared infrastructure, according to the company's incident write-up. That incident does not disprove an annual aggregate above 99%; it shows that an aggregate figure can coexist with meaningful degradation for affected users.

A separate number on OpenAI's Scale Tier page should not be confused with the keynote claim: the paid offering specifies a 99.9% uptime service-level agreement for Scale Tier traffic. That contractual figure has a defined customer tier and is not evidence that all Responses API traffic receives the same guarantee.

The API is becoming a platform for multi-step software, where a single customer action can generate many requests and failures in one step can undermine the entire workflow. OpenAI's DevDay claim points to substantially more activity on that infrastructure, but does not establish how much end-user demand it represents or how often production tasks complete without interruption. Those are the measures developers need when deciding whether to build critical workflows on the service. The growth rate and reliability claim are company statements; the keynote did not provide the details needed to compare them with customer-level outcomes.

Reader comments

Conversation for this story loads after sign-in.