Anthropic launches Claude Opus 5.5 with lower prices and a top-three benchmark score

Artificial Analysis ranks the high-effort fallback system third at 54, while Anthropic cut per-token prices 20% from Opus 5.

By · Published

Primary source: X - @synthwavedd

Why it matters

Opus 5.5 pushes the frontier-model contest toward cost per completed job, where token use and agent steps matter alongside benchmark scores. Its fallback-assisted result also shows why developers must inspect configurations before treating one leaderboard number as a clean model comparison.

A sleek, glowing digital display abstractly visualizes an intelligence index score of 58 with an upward-moving light effect in a modern research lab.

Anthropic launched Claude Opus 5.5 on September 22nd, pairing its newest frontier model with lower token prices, faster output and an independently measured Artificial Analysis Intelligence Index score of 54.

The launch is the first major model release since Anthropic co-founder and CEO Dario Amodei (@DarioAmodei) called for labs to slow the rate of capability advances so safety work could keep pace. Anthropic's answer is a model that advances the frontier while placing more weight on external evaluation, automated behavioral testing and safeguards that route some sensitive requests to older models.

Artificial Analysis puts the deployed system third

Artificial Analysis's model page gives Claude Opus 5.5 a score of 54 on its Intelligence Index, ranking it third among the 661 models listed on the tracker. The evaluator labels the tested configuration "Claude Opus 5.5 (Adaptive Reasoning, High Effort, Default Fallback)."

That direct listing supersedes the leaderboard image shared by leo (@synthwavedd) on X, which showed a 57.6 score and described the configuration as max effort. Artificial Analysis says the current index version combines 10 evaluations spanning agentic knowledge work, business workflows, coding, reasoning, knowledge reliability and long-context performance.

The configuration matters. Anthropic says its own evaluations used production safeguards. When those safeguards intervened, cybersecurity tasks were completed by Claude Opus 4.8, while biology and frontier-model-development tasks fell back to Claude Opus 5. The resulting score measures Anthropic's deployed system, including its routing behavior, rather than an isolated Opus 5.5 model answering every prompt.

Anthropic is selling efficiency with the capability jump

Anthropic says Opus 5.5 performs at roughly the level of Fable 5.1 on most work while costing 40% less to operate than Opus 5 on typical workloads. The per-token cuts are smaller: input drops from $5 to $4 per million tokens and output from $25 to $20. Cache reads fall from $0.50 to $0.20, a larger reduction for coding agents that repeatedly reuse the same repository context and instructions.

Anthropic also claims Opus 5.5 generates output over 30% faster than Opus 5. A separate fast mode, available through Claude Code and the Claude Platform, offers up to 2.5 times the speed at $8 per million input tokens and $40 per million output tokens.

The pricing shows where Anthropic expects the competition to move. Frontier models are increasingly sold as long-running workers, making the number of steps, tool calls and generated tokens as important as the model's nominal API rate. An agent that reaches the answer in half the turns can be cheaper even when its token price looks high.

Anthropic's own benchmark table gives Opus 5.5 a 66.4% result on Terminal-Bench 4.0, against 57.9% for GPT-6 Astra and 55.8% for Fable 5.1. It scored 54.4% on FrontierCode v1.1, narrowly above Astra's 53.3%, and reached 1,846 Elo on GDPval-AA v2.1, ahead of Fable 5.1 at 1,735. Those are vendor-published results, and Anthropic notes that configurations differ on some tests.

Astra retained the lead in parts of Anthropic's table. OpenAI's model scored 64.6% on Terminal-Bench-Science, compared with 58.7% for Opus 5.5, and edged it on Zapier's AutomationBench. The evidence supports a strong coding and knowledge-work launch, rather than a universal win across every workload.

Pacing still includes shipping

In his September essay, Amodei argued for "pacing the frontier", defining the policy as balancing capability development with stronger safeguards and third-party verification. He explicitly stopped short of calling for a halt to training or technical progress.

Anthropic says Frontier Design and METR evaluated Opus 5.5 before release. The company also says the model produced its best result to date on an internal automated behavioral audit and showed stronger resistance to prompt injection than Opus 5. These remain Anthropic's claims, backed by evaluations selected and reported as part of the launch.

That makes Opus 5.5 an early test of Amodei's proposed bargain: Anthropic intends to keep advancing models while arguing that external access, fallback systems and tighter deployment controls constitute meaningful restraint. The capability chart will get the immediate attention. The harder measure is whether those controls continue to work when developers give the model longer tasks, broader tool access and less human supervision.

Reader comments

Conversation for this story loads after sign-in.