Open-weight models surge past closed rivals in Vercel token traffic

DeepSeek V4 Flash and Step 3.7 Flash lead the shift, though Vercel's data covers its own gateway rather than the entire AI market.

By · Published

Primary source: Guillermo Rauch on X

Why it matters

Open-weight models now carry most token volume on Vercel's gateway, led by efficiency-focused releases from DeepSeek and StepFun. The split strengthens Vercel's bet that routing and observability will remain valuable as developers spread workloads across competing labs.

Open-weight models surge past closed rivals in Vercel token traffic — DeepSeek V4 Flash and Step 3.7 Flash lead the shift, though Vercel's data covers its own gateway rather than the entire AI market.

Vercel founder and CEO Guillermo Rauch (@rauchg) drew attention Tuesday to a crossover in his company's AI traffic: models with downloadable weights accounted for 53.9% of token volume routed through Vercel AI Gateway on August 11th, compared with 46.1% for closed-weight models.

"Keep an eye on this view," Rauch wrote on X, linking to Vercel's model leaderboards and a chart covering June 13th through August 11th. The data shows open-weight models gaining share during the two-month period and moving ahead of closed models near the end of the chart.

Rauch, who co-created Next.js before building Vercel around developer infrastructure, has positioned Vercel as a neutral layer between AI applications and model companies. The crossover supports that strategy: when developers regularly change models and providers, the routing layer becomes a durable point of control.

Vercel AI Gateway chart showing open-weight models at 53.9% of token volume
Open-weight models led closed models in Vercel AI Gateway token volume on August 11th. Graphic: Vercel.

DeepSeek and StepFun lead token volume

DeepSeek V4 Flash ranked first in Vercel's token-volume table with a 19.3% share. Step 3.7 Flash, whose weights StepFun publishes for download, placed second at 12.8%.

The highest-ranked closed model, Claude Opus 4.8, held 6.6%. DeepSeek V4 Flash 0731 and Claude Sonnet 5 each followed at 5.9%, while GPT 5.6 Luna accounted for 5.5%.

Vercel defines open-weight models as those that publish their weights for download. That classification should not be read as a broader judgment about whether a model is fully open source. Licenses, training data, source code, and usage restrictions can remain closed even when weights are available.

The figures also describe traffic on one intermediary. Vercel says its leaderboards aggregate anonymized production traffic routed through AI Gateway and update daily. They do not include requests sent directly to model companies or through competing inference platforms, and Vercel does not publish the absolute token count behind each daily share.

Vercel says the rankings are calculated across trillions of tokens. In July, Vercel opened the underlying leaderboard data under a Creative Commons license and added CSV downloads and a public export endpoint, making the daily shares available for outside analysis.

Tokens, requests, and spending tell different stories

Open-weight leadership in token volume does not translate into leadership across every measure. Vercel's request leaderboard was headed by GPT-5 nano at 15.1%, followed by OpenAI's text-embedding-3-small at 13.7%. DeepSeek V4 Flash ranked third with 5.8% of requests.

Spending was concentrated in higher-priced models. Claude Opus 4.8 accounted for 16.5% of AI Gateway spend, and Claude Opus 5 held another 14.7%. Kimi K3 placed third at 7.6%.

The split suggests developers are sending large token workloads through efficient open-weight models while reserving costlier closed models for a smaller set of requests. Token share can be driven by long contexts, verbose output, batch processing, or agent loops, so it cannot establish how many applications or customers have adopted a model.

It does establish that open-weight releases have become production infrastructure inside Vercel's customer base rather than an experiment at the edge of it. The leading models in the token table are efficiency-focused offerings built for high-volume workloads, where small differences in per-token cost compound quickly.

Vercel benefits from model churn

Vercel made AI Gateway generally available in August 2025. The service offers one API for hundreds of models, with usage reporting, provider routing, automatic failover and token prices tied to provider rates.

That model-agnostic pitch gets stronger as traffic fragments. Developers switching among DeepSeek, StepFun, Anthropic, OpenAI and other labs create demand for a common billing, observability and routing layer. Vercel can serve that role without predicting which model company wins.

Rauch's chart makes the strategic point with Vercel's own traffic. Open-weight models have crossed 50% of token volume on AI Gateway, while closed providers still capture much of the spending. Vercel sits in the path of both.

Reader comments

Conversation for this story loads after sign-in.