Anthropic says 354 Claude protein designs bound in lab tests

Anthropic released the prompts and data behind 1,440 designs, though none were tested for biological function or structurally solved.

By · Published · Updated

Primary source: Anthropic

Why it matters

Anthropic is testing whether a general frontier agent can operate specialist scientific workflows. The open release gives labs a reproducible protocol, while the missing functional evidence keeps the drug-discovery implications in check.

A stylized illustration shows an alpha helix protein model hovering over multiple lab wells on a white surface with numerical data.

Dario Amodei and Daniela Amodei's Anthropic sent Claude into protein engineering, where its agents ran autonomous design campaigns that yielded 354 proteins with measured binding across 15 interpretable targets.

Anthropic described the results in a post on X on Tuesday and published a 29-page technical report. Anthropic also released the prompts, predicted structures, provenance records and experimental measurements for all 1,440 designs on Hugging Face.

The work brings Amodei back toward his scientific roots. Before Google Brain, OpenAI and Anthropic, he earned a Princeton physics PhD studying neural circuits and worked at Stanford on cellular proteomes and cancer biomarkers. Daniela Amodei brought an operations and safety background from Stripe and OpenAI when the siblings founded Anthropic with former OpenAI colleagues in 2021.

Their founding thesis centered on building safety into frontier models from the beginning. The protein project extends that thesis into a practical question: can a general-purpose AI agent absorb a specialist's working knowledge, operate scientific software and produce physical candidates worth testing?

Amir Shanehsazzadeh, an Anthropic researcher and the report's sole listed author, designed the experiment around that question. He wrote a roughly 16,000-word protocol prompt covering the scientific and operational knowledge needed to run a binder campaign. Claude handled the individual design decisions.

What Claude actually did

Anthropic ran Claude Opus 4.8 and Mythos Preview as agents against 16 protein targets. The models researched each target, selected regions and binding sites, installed open-source design and structure-prediction software, generated candidates, optimized them computationally and returned 30 ranked designs per target.

Claude received no predetermined epitope, scaffold or protein sequence. Anthropic selected the targets, supplied cloud GPU accounts, set the budgets and time limits, approved non-scientific access requests and ordered the resulting proteins for synthesis. Humans interpreted the experimental results after the campaigns ended.

Single-target runs received 24 hours and a $10,000 compute budget. Multi-target runs received 48 hours and $50,000. Claude had to install and validate the scientific tools itself, then manage teams of sub-agents while keeping the work inside those constraints.

That distinction defines the technical contribution. Anthropic introduced no specialized protein-generation model. Claude coordinated existing open-source systems, including protein backbone generators, sequence-design tools and structure predictors. The report is evidence for frontier agents as scientific operators, rather than evidence that Anthropic has built a drug-design platform.

The same orchestration strategy runs through Anthropic's broader product push. Earlier Tuesday, RuntimeWire reported that Anthropic opened Claude Cowork to every paid plan, extending an agent designed to work across files and applications. Claude Science applies the underlying idea to research environments with specialist software, cloud compute and auditable artifacts.

The 27% comes with a denominator

Adaptyv Bio and Twist Bioscience, two paid contract research organizations, synthesized the designs and measured their binding. One target, mature GDF-8, aggregated during testing, leaving 1,320 designs across 15 targets with interpretable measurements.

Of those, 354 bound, for a reported hit rate of 26.8%. Claude produced at least one binder against 14 of the 15 interpretable targets. Forty-nine percent of the designs ranked first in their respective target campaigns bound.

Performance varied sharply by target. Claude produced 72 binders from 90 designs against TREM2 and 54 from 90 against VEGF-A. It produced three weak binders against the synthetic beta-barrel BBF-14 and none from 90 designs against maltose-binding protein. The computational confidence scores gave little warning that those campaigns would fail.

Mythos Preview's 24-hour single-target campaigns recorded a 35.1% hit rate, compared with 26.7% for its multi-target campaign and 22.6% for the Opus 4.8 multi-target run. The single-target format also received 2.8 times as much compute per target, so the experiment cannot isolate whether focus, budget or model behavior caused the difference.

The strongest comparison came on RBX1. Twenty-eight of Claude's 90 designs bound, compared with nine of 245 entries in a recent open competition. Anthropic's strongest reported RBX1 design produced an apparent dissociation constant of 3.9 nanomolar when tested on the same plate as the competition winner, which measured 45 nanomolar.

That result deserves attention, though it is not a clean human-versus-agent trial. Anthropic ran no matched campaign in which human experts received the same tools, compute and time. Results from four of the six open competitions used as comparisons were available to Claude during the design process.

The wet lab still gets the final vote

The report's sharpest limitation is biological. Anthropic measured binding, an early checkpoint in protein engineering. None of the designed proteins was tested for biological activity, and no protein-target complex was structurally solved. Every binding pose in the report remains a computational prediction.

The headline hit rate also counts sequence variants of the same underlying protein backbone as separate designs. When Anthropic counted only the best-ranked sequence from each of 809 generated backbones, 200 bound, producing a 24.7% rate. Affinity measurements for five oligomeric targets were apparent values, and each combination of model, campaign format and target was generally run once.

Those caveats put distance between this experiment and a therapeutic program. Companies such as Latent Labs, Manifold Bio and Generate:Biomedicines are building specialized models, screening systems or drug pipelines around protein engineering. Google DeepMind's AlphaProteo has paired binder generation with structural and functional validation on selected targets.

Anthropic's immediate opportunity sits one layer above those systems. Claude can choose among scientific tools, install them, manage compute and preserve a record of its decisions. A laboratory could adapt the released protocol without building its own agent framework from scratch.

Publishing the complete design set makes the release more useful than a table of successful examples. The Hugging Face repository includes failed campaigns, lower-ranked candidates, raw binding measurements and the records needed to inspect which tools produced each design. The data is licensed under CC BY 4.0, while the released scripts use the MIT license.

Shanehsazzadeh also used Claude Science for the post-campaign analysis, figures and manuscript drafting, according to the report. He checked the analyses against the released data and accepted responsibility for the paper. Anthropic has effectively published an AI-run experiment and an AI-assisted account of that experiment, with enough underlying material for outside researchers to challenge both.

Claude did not produce a drug. It showed that a general agent could run most of a computational protein-design campaign, deliver physical candidates that bound in independent testing and expose its work for replication. The remaining steps - structure, function, safety and efficacy - are where biology stops behaving like a software benchmark.

Reader comments

Conversation for this story loads after sign-in.