On August 18, 2026, Anthropic published results from two wet-lab collaborations testing whether Claude can do experimental science end to end rather than merely advise on it. In the first, run with Adaptyv Bio and Twist Bioscience, Claude was asked to design protein binders - small proteins engineered to latch onto a target protein - a task Anthropic describes as having “historically taken protein engineers months of computation, optimization, and screening per target.” Across 1,320 designs, the campaign produced 354 confirmed binders and succeeded against 14 of 15 targets.
The reported hit rates were 26.7 percent for a Mythos Preview model and 22.6 percent for Opus 4.8 in multi-target mode, rising to 35.1 percent for Mythos Preview when working a single target at a time. Anthropic contrasts this with an industry norm it puts at 10 to 15 percent. The company says Claude produced high-affinity binders for at least six targets and matched or beat the best published affinities on at least four, including a 40 percent hit rate against RBX1 where human competitors in the same setting reached 3.7 percent. Fifteen of the designs used beta sheets, a structural motif that generative binder methods usually avoid.
The failures are reported alongside the successes, which is what makes the writeup useful. Claude produced zero confirmed binders against maltose-binding protein from 90 designs, and only three modest-affinity binders against BBF-14. Anthropic also notes it does not know why Opus 4.8 succeeded on TNF-alpha while Mythos Preview did not. The second collaboration was narrower and more mundane: Claude Opus 5 was handed raw NMR and LC-MS spectroscopy files and asked to characterize compound identity and purity. It finished the NMR analysis in 23 minutes and the LC-MS in 19, landed hydrogen counts within 0.08 units of the contract lab’s, reported 96.4 percent purity against the lab’s 96.33 percent, and reverse-engineered an undocumented vendor file format along the way.
For a business or technical leader, the analytical chemistry result is arguably the more immediately actionable of the two. Binder design is a headline capability with an unexplained failure mode and a hit rate reported by the vendor rather than an independent evaluator; spectroscopy interpretation is routine, high-volume, expensive labor in every chemistry organization, and the reported accuracy is checkable against an existing lab process. The honest read is that these are vendor-reported results from a small number of campaigns, not peer-reviewed replication, and the sample of 15 targets is thin. But the direction of travel matters: the constraint on applying frontier models in the lab is shifting from model reasoning to the cost and turnaround of physically testing what the model proposes.