Anthropic loosens Claude Fable 5's biology safeguards to cut false positives

On August 7, 2026 Anthropic published an update to the biology safeguards on Claude Fable 5, aimed squarely at false positives rather than at loosening the underlying threat model. Anthropic protects against biological misuse in part with safety classifiers, smaller automated systems that detect when the model is asked to perform a safeguarded biology task or to produce a harmful output. When a classifier fires, the product falls back to a less capable model. That fallback behavior was catching a large volume of ordinary health and education questions.

The change was made by rewriting the classifier’s constitution, the instruction set that defines what the classifier should treat as safeguarded. Anthropic says it rewrote that document over several weeks, solicited feedback from internal and external experts, generated updated training data from the revised constitution, and retrained the classifier. The effect is a repositioned boundary rather than a removed one: many more benign requests pass, while requests tied to dual-use professional research continue to be blocked.

The reported reduction in biology-related fallbacks is approximately 85 percent across Anthropic’s product surfaces. Reductions in total fallbacks vary sharply by surface: 67 percent on Claude.ai, 55 percent on Cowork, 17 percent on Claude Code, and 7 percent on the Claude Platform. Anthropic cites interpreting lab results, understanding symptoms, and learning biology in an educational context as the cases users should notice, and says healthcare professionals will be able to get more support on clinical tasks.

The business relevance is that false positive rates in safety systems are now a published, quantified product metric at a frontier lab, not an internal quality complaint. Over-blocking has a real cost in regulated and clinical settings, where a silent downgrade to a weaker model is worse than a refusal because the user cannot tell it happened. Anthropic’s framing, that a classifier constitution can be rewritten and retrained against expert review without moving the underlying capability threshold, is the operational template other deployers of guardrail classifiers will be measured against.

Sources

Last verified August 17, 2026