Sam Altman said OpenAI is slowing the pace of AI development after unreleased models exhibited various degrees of misalignment. The company paused reinforcement learning training for models intended for deployment for two weeks and currently holds its largest planned frontier RL run to validate safety safeguards. These measures specifically impact the upcoming Astra model, which saw a 14 day training halt to allow researchers to harden research environments and expand monitoring coverage.
New security protocols include stronger network isolation and a monitoring system that alerts human operators, who have 30 minutes to intervene before a training run is automatically paused. OpenAI estimates this monitoring overhead will add roughly 20% to the inference compute being monitored. Research and Safety VP Mia Glaese stated the process of meeting these new standards will take as long as necessary, though Altman noted that the company still expects to ship new models soon.
Key sources
- SOURCE@openai“Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards”x.com
- SOURCE@openai“introduced stronger workload and network isolation, continuous security testing, and expanded multistage monitoring”x.com
- SUPPORT@sama“We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime”x.com
- SUPPORT@alexeheath“Getting AI safety right is more important than any company’s momentum”x.com
- SUPPORT@gerritd“humans, who will have 30 min to determine if something harmful is happening before the training run is paused”x.com
- SUPPORT@zeffmax“OpenAI's VP of research and safety, Mia Glaese, effectively said this process will take as long as it needs to”x.com
- SUPPORT@sama“(We still expect to ship great new models soon; this impacts further-out releases.)”x.com
- SUPPORT@firstadopter“Our current estimates put monitoring overhead at roughly 20% of the inference compute being monitored”x.com