Arena ranks GPT-6 Astra Max first for web development, Claude Fable 5.1 Max first for agents

Arena's founders are turning model evaluation into a business by dividing "best" into tasks, costs and confidence intervals.

By · Published

Primary source: Arena

Why it matters

Arena is becoming part of the purchasing layer for AI models, using human votes and workflow traces to influence which systems developers adopt. Its latest tables also show why no leaderboard should be read as a universal measure: GPT-6 Astra Max led WebDev, while Claude Fable 5.1 Max led Agent Arena and GPT-6 Astra Max ranked second.

A person's hands interact with a tablet displaying abstract data visualizations that compare the performance of different AI models for specific tasks.

Arena, the AI evaluation platform co-founded by Anastasios Angelopoulos (@ml_angelopoulos), Wei-Lin Chiang (@infwinston) and Ion Stoica (@istoica05), ranked OpenAI's GPT-6 Astra Max first in its September 11th WebDev snapshot. Anthropic's Claude Fable 5.1 Max finished second, while Claude Opus 5 Max took third.

The order changes when the job changes. Arena's separate Agent Arena leaderboard, dated September 9th, put Claude Fable 5.1 Max first and GPT-6 Astra Max second on agentic tasks involving tools. The split is the useful result: model quality is becoming specific to a workflow, and Arena is building its business around measuring those differences before buyers commit engineering time and inference budgets.

Angelopoulos and Chiang started Chatbot Arena as a Berkeley research project in 2023. Anonymous models answered the same prompt, users picked the response they preferred, and the votes fed a public ranking.

The founders brought complementary versions of the same evaluation problem. Angelopoulos studied reliable AI and statistical guarantees at Berkeley after earning an electrical engineering degree at Stanford, and his doctoral advisers included Michael Jordan and Jitendra Malik. Chiang worked on Vicuna, FastChat and SkyPilot inside Berkeley's systems research community. Stoica, a Berkeley professor, previously co-founded Databricks, Anyscale and Conviva.

Their side project became Arena Intelligence in 2025. By September 2026, Arena was no longer measuring chat answers alone. It was recording how models plan, edit files, run commands, debug applications and arrive at a rendered result.

GPT-6 Astra Max won this task, at this price

The WebDev table covered 128 models and 679,295 votes. GPT-6 Astra Max received a score of 1800 from 2,281 votes, with a 16-point confidence band on either side. Claude Fable 5.1 Max scored 1758 from 3,036 votes, followed by Claude Opus 5 Max at 1687 from 12,087 votes.

Arena listed GPT-6 Astra Max at $10 per million input tokens and $50 per million output tokens, the same rates shown for Claude Fable 5.1 Max. Claude Opus 5 Max was listed at $5 and $25, respectively.

Those prices make the rest of the table as important as the podium. Alibaba's Qwen 3.8 Max 0902 ranked fourth at a listed $2 per million input tokens and $6 per million output tokens, although Arena marked the result preliminary. Alibaba's Qwen 3.8 Flash Next placed ninth at $0.16 and $0.47. Buyers choosing a model for production still have to decide how much additional ranking performance is worth paying for.

Arena's Code Arena methodology evaluates models in a controlled environment where they plan, generate, edit and debug applications while Arena records tool actions and rendered results. That design gets closer to how an AI coding system behaves during a build than a one-shot code-generation test.

It remains a controlled, preference-based environment. The WebDev snapshot does not establish how much code a model can ship inside a production repository, how well it works with a particular framework, or whether it reduces engineering time after review and maintenance are counted. The task mix, user population and voting behavior all shape the result.

Vote counts also vary sharply. GPT-6 Astra Max reached first place with fewer than one-fifth as many votes as third-ranked Claude Opus 5 Max. Arena publishes confidence intervals and rank spreads to make that uncertainty visible, and several recent entries carry preliminary labels.

The founders are selling measurement, not a permanent winner

Arena has faced scrutiny over how voting-based rankings can be influenced. A 2025 paper, The Leaderboard Illusion, argued that uneven model sampling and private testing could advantage large providers. Arena disputed parts of the analysis and outlined additional disclosure and testing policies.

A separate peer-reviewed study, co-authored by Angelopoulos, found that voting-based leaderboards could be manipulated when defenses were absent. The researchers worked with Arena on mitigations including rate limits, malicious-user detection, bot protection and login controls. The episode established an unavoidable constraint on Arena's model: human feedback becomes valuable at scale, while the same openness creates a surface for strategic voting and selection effects.

Arena's response has been to collect richer evidence around each result. Code Arena records tool actions and rendered outputs. Agent Arena reports task completion, steerability, tool hallucination, median task cost and output-token use. A single ordinal ranking remains the easiest product to read, but the underlying business increasingly depends on selling a detailed account of how a model behaves.

That business has attracted $250 million in disclosed venture funding. Arena raised a $150 million Series A on January 6th, led by Felicis and UC Investments, at a reported $1.7 billion post-money valuation. Andreessen Horowitz, The House Fund, LDV Partners, Kleiner Perkins, Lightspeed Venture Partners and Laude Ventures also participated.

Arena said on June 29th that its enterprise evaluation service had reached a $100 million annualized revenue run rate eight months after launch. Arena also claimed more than 10 million monthly visitors, 700 million conversations and 82 million votes. Those operating figures are company-reported.

The commercial loop is straightforward. Arena gives users access to competing models, users generate prompts and judgments, and Arena packages that real-world feedback for model laboratories and enterprises. Each new evaluation category expands the set of decisions Arena can influence, from choosing a chatbot to selecting a coding model or an agent that can safely operate tools.

The WebDev result supports Arena's task-specific approach. GPT-6 Astra Max led Arena's front-end development table on September 11th. Claude Fable 5.1 Max led Agent Arena two days earlier, with GPT-6 Astra Max in second place. Arena's founders are betting that differences among those rankings will make independent evaluation infrastructure more valuable than any individual crown.

Reader comments

Conversation for this story loads after sign-in.