Anthropic’s Claude Is Now Designing Proteins Like an Expert

Anthropic let Claude run whole protein design campaigns on its own, and the wet-lab numbers beat typical human-led hit rates. Here is what happened.

Anthropic’s Claude Is Now Designing Proteins Like an Expert

Anthropic just published the most complete demonstration yet of an AI agent running a protein design campaign from start to finish. Claude researched each target, picked where to bind, installed its own tools, and delivered ranked designs for the lab. Humans approved access requests, kept the infrastructure running, and placed the synthesis orders.

A minibinder is a small protein engineered to latch onto a specific target, and that latching step is how a large share of medicines work. Designing one from scratch has historically taken a specialist weeks to months per target, and the machine learning shortcuts still demand days of expert orchestration.

The wet-lab results are the headline. Of 1,320 designs that two contract labs synthesized and tested, 354 bound their target. That is a 26.8% hit rate overall. Anthropic puts the typical range for protein design campaigns at 10 to 15%.

The timing is awkward for Anthropic. This is the same company whose models were caught hacking three companies during security tests.

That makes this the constructive flip side of a rough year for agent behavior, from agents escaping containment at OpenAI to the runaway-objective cases METR keeps documenting. An agent let loose on a hard scientific problem, with every result verified by outsiders, is the version of autonomy worth having.

What Claude actually did

Anthropic didn’t build a new biology model. Claude orchestrated open-source tools the field already uses, installing programs like RFdiffusion and BindCraft from their public repositories and combining them across 24 different workflows.

A protocol prompt of roughly 16,000 words told it how to run a campaign, and it left every design decision to the model. The campaigns ran 24 to 48 hours each, with up to 12,500 H100 GPU hours for the multi-target run.

Claude chose the epitopes, generated and optimized candidates, screened for ones that would express and stay soluble, and returned 30 ranked designs per target.

The human contribution was thin by design. People picked the targets, wrote the prompt, ordered the synthesis, and interpreted the measurement data. In between, the team only had to nudge the system back awake after infrastructure hiccups.

The protein design numbers that matter

Mythos Preview and Opus 4.8 hit overall rates of 26.7% and 22.6% when designing against all 15 targets at once in 48-hour sessions. Focusing one target at a time lifted Mythos Preview to 35.1%, though with 2.8 times the compute budget per target. The report doesn’t separate the higher hit rate from the compute budget.

The target-level results are less tidy. On TNFα, the target behind drugs like Humira, only Opus 4.8 succeeded, producing 12 binders from 150 designs after several expert groups reported zero hits.

Cross-species binding was a secondary goal, yet 130 of 233 tested binders also grabbed the mouse version of their target, which matters for animal studies.

On the immune receptor TREM2, 72 of 90 designs bound. Adaptyv Bio’s competition on that target topped out at 38.3%. The clearest flex is RBX1: in an open design contest, 9 of 245 human entries bound, while Claude landed 28 of 90.

Anthropic had the contest’s winning RBX1 design rebuilt and measured on the same assay plate as Claude’s best. The winner bound at 45 nM. Claude’s bound at 3.9 nM, roughly ten times tighter.

Adaptyv Bio, one of the two validating labs, says the designs matched or surpassed expert designers on hit rate and affinity, and would have won 5 of 6 of its protein design competitions. Both labs tested the designs blind, not knowing which model designed what.

Why any lab can copy this

Every tool Claude used is open source, which is the quiet headline. Anthropic published the prompts, all 1,440 computational designs, and the full binding data on Hugging Face under CC BY 4.0.

The release includes raw sensorgrams, per-design provenance, and co-folded structure predictions, so any lab can rerun the protein design stack or use it as a benchmark.

The post also includes a result that working chemists will care about. Given a contract lab’s raw instrument files and a two-sentence prompt, Opus 5 returned finished NMR and LC-MS analyses in 23 and 19 minutes. Its purity estimate of 96.4% matched the lab’s own 96.33%.

Don’t expect to reproduce this in the Claude app tomorrow. Anthropic notes that life science research tasks are currently blocked in its most capable model, and says an access program for scientists is a priority.

The limits, and they are real

Anthropic and The Decoder are careful about where this protein design result stops. There wasn’t a parallel human campaign as a control, so the results don’t prove the system is better than humans. Only binding was measured, not structure or biological effect, and no design was structurally resolved.

The failures are instructive too. Against MBP, a bacterial protein with a smooth surface, zero of 90 designs bound. The confidence scores did not flag that target or the weak BBF-14 results ahead of time. Independent review of the whole campaign is still pending.

The pattern I keep watching this year: agents stopped being a chat feature and became labor. When that labor gets scoped to a wet lab and verified by people who don’t care about hype, you get numbers you can actually trust.

The scary version of this story is an agent with a browser and no oversight. The useful version is a tireless research colleague that never sleeps, and it just had its best week yet.

Tony Simons

Reviewed & Written By

Tony Simons

Independent tech reviewer and creator of Tony Reviews Things. 14 years of hands-on testing, software auditing, and workflow automation. I test the gear so you don't waste your money on junk.

Submit a Take

Your email address will not be published. Required fields are marked *