OpenAI pauses frontier RL training and says it will evolve its Preparedness Framework

On August 18, 2026 OpenAI published “Pacing model development in an era of cyber-critical capabilities,” a company post disclosing that it had slowed its own scaling. It says two things forced the change: the OpenAI-Hugging Face incident, and preliminary evidence that an upcoming model family called Astra “may meet the Critical cybersecurity capability threshold under our Preparedness Framework.” The concrete action was a two-week pause in reinforcement learning training on the latest models intended for deployment, while OpenAI hardened and red-teamed its research environments. The status to be precise about: the pause is over, but the post states that “Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.” That is a self-imposed hold under a voluntary framework, not a regulatory order, and nothing here was compelled by any government.

The post describes three reinforcing safeguards - monitoring, alignment, and security - and details what changed in each. On security, OpenAI says it paused frontier model inference in research clusters for runs that could execute code or reach the internet immediately after the Hugging Face incident, then restored workloads individually under new controls covering workload sandboxing, network isolation, removal of shared services, reduced standing privileges, and continuous automated attack simulation. It states that Astra and other cyber workloads now require the strictest security tier, and that “a significant number of workloads remain paused until they are fully migrated,” with safety and alignment work prioritized for migration first.

The monitoring disclosure is the most operationally specific part. OpenAI describes a multistage system starting with activation classifiers that inspect the model’s internal activity at every sampled token and escalate to higher-compute automated investigators, aiming to raise an alert within 30 minutes of concerning activity; if safety, security and research teams cannot rule out a false positive within another 30 minutes, they are expected to pause the activity. This monitoring is required for all RL training and evaluations involving tools at “Sol capability or higher,” and, after the August 7 determination about Astra, for all Astra inference with tools. OpenAI puts the cost at “roughly 20% of the inference compute being monitored,” while noting the figure varies substantially by workload. On the framework itself, the company commits only to evolving the Preparedness Framework to span training as well as deployment; no new version, threshold, or date is given.

For a leader, the useful signal is not the safety language but the admission of an economic tradeoff made public. A frontier lab has now said on the record that its capability progress outran its internal controls, that fixing that “incurred great cost and delays to frontier research,” and that continuous monitoring carries a roughly one-fifth compute tax on the workloads it covers. That number is a planning input for anyone budgeting agentic AI at scale, and it is a benchmark regulators and customers can point at when a vendor claims monitoring is free. The limits are equally plain: every claim here is OpenAI’s own, unaudited by any external party; the Astra assessment is described as preliminary rather than concluded; and the commitment to change the Preparedness Framework is an intention with no published deadline or content. Voluntary self-restraint that a company can lift on its own schedule is evidence of judgment, not of accountability.

Sources

Last verified August 24, 2026