On September 22, 2026, the day Anthropic released Claude Opus 5.5, METR published a summary of its independent predeployment evaluation of the model. METR had API access for 10 business days and focused on one question: how far the model could accelerate or automate AI research and development. It ran five capability tasks: a budget-constrained variant of the NanoGPT speedrun, a conceptual argumentation task about language models, “Train a Program” (replicating a program’s behaviour through machine learning training), a gaming-bot task in which the model writes programs that play video games through an API, and “Sunlight”, an open-ended research and report-writing task. It supplemented these with capability trend analysis, questionnaires and interviews with Anthropic researchers.
METR’s conclusion is that acceleration from this model would be slightly higher than for Claude Fable 5.1, but that the model is unlikely to be able to fully automate AI R&D. It found incremental improvement over its predecessor on both measurable and harder-to-verify tasks, with remaining qualitative weaknesses in long-horizon reasoning and research judgment, and said the model is still likely to noticeably accelerate researchers and automate limited parts of R&D. METR quotes a preliminary AI R&D report estimating about 1.5x overall acceleration in capabilities due to AI, with perhaps a 30 percent chance of 2x, while noting that the report did not say which time period the estimate covers.
Why it matters: automated AI R&D is the threshold several frontier safety frameworks treat as the trigger for much stricter controls, and an outside evaluator publishing a reasoned judgment on launch day is a check that does not depend on the developer’s own word.
What it does not show: ten business days is short, METR says some of the information it relied on cannot be disclosed, and the summary gives no new time-horizon figure for the model. “Unlikely to fully automate” is a judgment about a threshold, not a measurement of how much research the model actually speeds up.