Anthropic spotlights Opus 5.5 demos, with a human still in the loop

Three days after launch, Claude's account shared user experiments; one interactive lens lesson came from a designer using his own project files and a maximum-effort run.

By · Published

Primary source: Claude

Why it matters

Anthropic's showcase makes Opus 5.5's coding pitch tangible, but the strongest example also shows the human inputs behind a polished result: domain experience, prior project files, and a deliberate effort setting.

A person's hands work on a computer screen displaying complex 3D design software in a modern studio environment.

Anthropic used a September 25th post from Claude's official account to spotlight experiments people had made with Claude Opus 5.5, just three days after the model's September 22nd launch. The examples are user-built demonstrations, not a new product release or an independent evaluation. One offers a useful look at what the showcase can prove: a model can help turn a specific brief into an interactive artifact, while the person directing it still supplies much of the context.

Ryan Sael's post on X

poster=/api/storage/public-objects/tweet-videos/anthropic-opus-5-5-user-demos-operator-poster-e7370f96.jpg|Video from @RyanSael on X

That example came from Ryan Sael (@RyanSael), whose interactive camera-lens lesson was also featured in RuntimeWire's earlier reporting on the build. The lesson lets visitors adjust focus and aperture and see how those choices affect the sharp area of an image. Sael said the run took 86 minutes and cost $25.66. Those are his figures for one run, not a general estimate for building a similar project.

The operator is part of the result

Sael's setup matters. He asked Opus 5.5 to draw on files from his earlier projects and used the model's maximum-effort setting. His portfolio includes interactive web work, alongside previous roles building VR and AR experiences at FXMedia Singapore, leading interface design at Bridestory, and designing charity software at N3O, as detailed in the earlier RuntimeWire report. His finished lesson is a genuine, inspectable artifact. The one-run framing still includes a practiced designer, a specific learning goal, prior materials, and a high-effort model setting.

The context helps explain how Sael got to a polished result. It shows where an agent can help a builder move from an idea to an interactive page, and why the result should not be read as a blank-prompt test. Visitors can manipulate the lens illustration themselves; that establishes what the page does, not whether every optical calculation behind it is correct.

Anthropic introduced Opus 5.5 as the first model in its Claude 5.5 family and says it costs 40% less to run than Opus 5 on typical workloads. The company's listed API rates are $4 per million input tokens and $20 per million output tokens, compared with $5 and $25 for Opus 5. Those prices provide context for Sael's reported bill, but without his token usage they do not independently confirm it.

The launch pitch centers on long-running coding and knowledge-work agents. Anthropic says Opus 5.5 can handle tasks such as codebase migrations, computer use, and producing documents and spreadsheets. A user showcase speaks to a different question: what a person can make when they guide the model toward a finished product. Anthropic's post links to several creators' experiments, but the linked examples do not all come with enough detail in the post itself to compare their methods or results on equal terms.

Demos meet a narrower benchmark

An independent check offers a more bounded comparison. Sonar's evaluation tested a pre-release build on its Java benchmark and reported an 87.68% pass rate for Opus 5.5 across 544 test-backed HumanEval and MBPP tasks, versus 88.6% for Opus 5. The broader run recorded 4,444 tasks, but only those 544 counted toward the pass-rate figure. Sonar also found the newer model wrote 27.5% less code and generated 42% fewer total findings in that test. The results are specific to Sonar's benchmark and pre-release build; they do not validate Anthropic's full set of performance or cost claims.

Sael's demo and Sonar's benchmark answer different questions: one documents a guided build; the other tests model output across a fixed set of Java tasks. A demo makes a capability tangible, but it also reflects choices made by the person who prompted, configured, and selected the work. A benchmark holds some of those variables steadier, while answering a narrower question. Neither substitutes for observing how an agent behaves in a team's actual codebase and review process.

Anthropic's post is a small but telling choice of launch follow-up: it asks users to show what they explored, rather than asking the market to rely only on benchmark tables. Sael's lens lesson puts a human operator clearly in the frame. His design experience and reuse of earlier files are part of how the model's work became a coherent teaching page. For people weighing agent workflows, the practical point is that the output can be impressive while the operator's judgment remains part of the system that produced it.

Reader comments

Conversation for this story loads after sign-in.