Robocurve says Opus 5.5 matches Astra on robots for 21% less

Across 360 real-robot trials, Opus 5.5 averaged 36% progress at an estimated 90 cents per run; Astra completed more tasks outright.

By · Published

Primary source: X

Why it matters

The benchmark puts a measurable cost and performance comparison on real robot trials, while its selection of Astra's easiest tasks and low completion counts show why average progress alone is an incomplete proxy for useful autonomy.

A sleek robotic arm performs a precise task with a small component on a workbench in a modern lab.

On September 23rd, Jay Chooi (@chooi_jeq) and four colleagues at Robocurve reported that Opus 5.5 achieved nearly the same average progress as GPT-6 Astra across 360 real-robot trials, at a lower estimated model-inference cost. The comparison comes from Robocurve's own RoboDojo-RC Tier 1 report, which also makes clear what the headline averages leave out: Astra completed five tasks, while Opus 5.5 completed one.

https://x.com/chooi_jeq/status/2102847198690210031

poster=/api/storage/public-objects/tweet-videos/robocurve-opus-55-robotics-benchmark-poster-48a82f43.jpg|Video from @chooi_jeq on X

The benchmark ran 120 trials per model across six tabletop tasks, with 20 trials per model on each task. Robocurve scored progress against task-specific rubrics, so its 36.0% mean for Opus 5.5 and 36.7% for Astra reflect partial progress as well as completed tasks. The older Opus 5 averaged 19.9% progress. Robocurve estimated costs at 90 cents per Opus 5.5 trial, $1.14 per Astra trial and $1.76 per Opus 5 trial.

Those costs need a qualification. Robocurve adjusted Astra's estimate to assume successful prompt caching; it used recorded cache usage for the two Opus models. The adjustment changed Astra's cost estimate, not its recorded results, and Robocurve says no new model responses were generated for the calculation. Its report says the cost figures exclude robot operation and are not invoice amounts. The 21% gap is therefore a comparison of estimated model-call costs under stated assumptions, not a full estimate of deploying a robot.

Chooi is a co-founder of Robocurve, a public-benefit corporation built around independent robotics evaluations. His path into the work runs through AI evaluation and policy: his biography says he was a research fellow at MATS, graduated from Harvard in 2026, and was elected a Rhodes Scholar in 2025. The report credits research leads Sravanthi M (@sravanthi6m) and Achu Menon (@achumenon_), alongside Sabrina Zou, Tzu Kit Chan and Chooi.

The findings favor Opus 5.5 over Opus 5 on five of six tasks, but they do not show a clean across-the-board win against Astra. Opus 5.5 scored higher than Astra on capping a pen and packing and pouring fruit. Astra led on stacking bowls and standing bottles upright; the models were close on putting objects in a safe and sorting objects. On the bottle task, all three models barely made progress: Opus 5.5 averaged 0.5%, Astra 2.5% and Opus 5 5%.

Completion counts add another layer. Astra fully completed five of its 120 trials, against one for Opus 5.5 and two for Opus 5. Robocurve's progress score gives credit for intermediate steps, useful when a robot gets partway through a task, but that measure is different from reliably finishing the job. The report's headline comparison is about average rubric-scored progress, not equal task-completion rates.

The test also has a deliberately limited scope. Robocurve says it first evaluated Astra on 18 RoboDojo tasks, then selected the six easiest by Astra's scores for this Tier 1 comparison. It adapted the setups, instructions and rubrics, so the resulting scores are not directly comparable with the original benchmark. Robocurve's human graders knew which model produced each run, a potential source of unconscious bias the report acknowledges. It also notes that objects were reset by hand, leaving room for placement variation.

All three models used the same Inspect Robots agent policy at medium effort, with a 40-call budget, 900-step cap and 25% speed cap, controlling bimanual robot arms. That creates a useful controlled comparison, while leaving open how rankings would change under different effort settings or control limits. Robocurve also cautions that matching the nominal effort setting does not guarantee equal reasoning budgets across model providers.

The practical contribution is the evaluation record. Robocurve published all 360 trial traces and the benchmark materials, while its Inspect Robots framework is open source. That lets other researchers inspect the runs and repeat the test rather than rely on a polished demonstration. For Chooi's company, publishing the limits alongside the result is central to the pitch: robotics progress should be measured through repeatable physical trials, even when the outcome is partial progress rather than a robot reliably completing the work.

Reader comments

Conversation for this story loads after sign-in.