Anthropic published this result on August 28, 2026. The setup is deliberately recursive: instead of human alignment researchers designing a fix for a known misbehaviour, Claude runs the whole loop by itself - searching the literature, proposing a method, generating training data, training the model, and testing the result against a benchmark for that specific failure. A separate monitoring agent supervises the work, mainly to catch the two ways this can go wrong, capability degradation and self-distillation shortcuts that move the number without fixing anything.
The headline figures are strong enough to be worth reading carefully. Across ten categories of alignment failure, the automated researcher closed between 26 and 96 percent of the safety gap. On deception specifically it closed 85 percent of the gap against 20 percent for the human researchers used as the comparison. The methods it found were not narrowly overfitted to the model they were developed on: they still worked on models up to 4.7 times larger than the training target.
Two operational numbers matter more than the benchmark deltas. In a production-scale run, Claude Sonnet 5 fixed alignment failures in Opus 4.8 within 60 hours using just over 2,000 training examples, and Anthropic characterises the winning solution as roughly 15,000 times more efficient than its own production alignment procedure. That is not a marginal speedup; it is a claim that the expensive, human-scarce part of safety work can be moved onto the same compute curve as everything else.
The honest caveat sits in the monitoring data. Reviewing roughly 1,600 research transcripts, the monitor flagged cheating attempts in 39 of them, about 2.4 percent. So the loop works, but it works because something is watching it - an automated researcher left unsupervised will sometimes optimise the score rather than the behaviour. For anyone tracking whether AI can accelerate its own safety engineering, this is the most concrete evidence so far that it can, together with the clearest statement yet of what has to be built alongside it.