OpenAI says GPT-6 Astra and Sol now stream 50% faster
Tibo Sottiaux put the change at 50 tokens per second, up from 30, for subscribers using OpenAI products and supported partner apps.
By Ryan Merket · Published
Primary source: X
Why it matters
OpenAI is extending a speed improvement into third-party coding tools that draw on ChatGPT subscriptions, making responsiveness part of its competition for the account relationship and recurring usage.

OpenAI says it has increased the default response speed of GPT-6 Astra and GPT-6.1 Sol by about 50% for subscribers using its products and partner apps. Thibault Sottiaux (@thsottiaux), OpenAI's head of core products, wrote in a thread on X on October 5th that the change would reach products using Sign in with ChatGPT without users changing their settings.
Sottiaux put the new output rate at roughly 50 tokens per second, compared with 30 before, and said the rollout should be felt within two hours of his post. Those figures do not line up exactly with the headline percentage: going from 30 to 50 tokens per second is about a 67% increase if both numbers are exact. Sottiaux described the rates as approximate, and did not specify how OpenAI measured them or whether they apply uniformly across products, plans, and tasks.
The comparison is about how quickly text streams, not how quickly a model completes a job. A coding agent can spend time planning, calling tools, running code, and checking its work; token output speed alone does not measure that full cycle. Sottiaux also said the models use a highly optimized tokenizer and need fewer tokens to complete tasks, but gave no token-consumption figure or benchmark alongside that claim. OpenAI has not supplied independent measurements for the new speed in the announcement.
The improvement reaches beyond ChatGPT and Codex through OpenAI's account and subscription connections with outside tools. Sottiaux named OpenCode, Pi, Amp, and Devin as examples of products using Sign in with ChatGPT. OpenAI describes the feature as a way for eligible users to apply their ChatGPT plan to AI requests in participating apps, with usage counted against their plan rather than requiring an API key. The company says users can set weekly limits for individual apps and review their usage in ChatGPT settings. That means a faster model can affect coding workflows running inside a separate developer product, while the underlying model access and usage limits remain tied to OpenAI's subscription.

For OpenAI, performance in these integrations is part of the product, not just a model specification. A developer may spend most of a work session in an outside coding tool, but if that tool draws on ChatGPT-plan usage, OpenAI still controls the account connection and the usage allocation. Sottiaux has framed OpenAI's product strategy around bringing agent capabilities to people beyond the technical audience. In an August interview with TechCrunch, he said he leads OpenAI's core products, including the API, agent infrastructure, enterprise products, ChatGPT, and Codex, and described the effort as making coding-agent capabilities available to a wider population.
OpenAI positions GPT-6.1 Sol as a lower-cost option for coding, computer use, and professional work, while GPT-6 Astra is its flagship model for demanding reasoning and coding tasks. Faster streaming can make either feel more responsive in interactive use, particularly when a person is watching an agent work. It does not establish that either model produces better answers or finishes a complex task sooner.
The distinction matters to subscription users because app activity can draw from the same overall plan allowance. OpenAI's current help material says users can set per-app weekly limits, and that an app's limit does not create a separate pool of usage. Faster generation could improve the experience while leaving the amount of work a customer can complete before hitting a usage limit unchanged. The announcement does not say whether OpenAI adjusted those limits, changed the models themselves, or altered the underlying compute allocated to each request.
Sottiaux's claim therefore sets a specific expectation for users: quicker text output, with no setup change. Whether the increase improves end-to-end coding-agent performance depends on latency and task-completion measures that the post did not provide.