OpenAI Unveils New Security Policies to Mitigate Risks in AI Model Testing
On Tuesday, OpenAI announced a set of new security policies aimed at enhancing safety measures during the testing of its AI models. The updated protocols include detailed monitoring throughout the development process and an increased focus on alignment and security in the post-training phase. The company stated in a blog post, “As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks.”
This announcement marks one of the first public adjustments to OpenAI’s safety practices since the Hugging Face incident, disclosed on July 26th. However, OpenAI representatives clarified that these measures are not solely a reaction to that incident; they were partly prompted by the cybersecurity features of the upcoming Astra model and the overall rapid advancement in AI technologies.
As part of the new policies, OpenAI shared that it temporarily halted reinforcement learning activities for two weeks following the Hugging Face incident, although it has since resumed training for less risky models. “Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the blog post noted.
Amelia Glaese, OpenAI’s VP of research, emphasized to reporters that the rigor of security controls will intensify as models become more advanced, with larger models facing the most scrutiny. “We have put in place requirements and expectations for safe development,” Glaese remarked, adding that “those requirements and expectations vary with the level of risk that we see.”
The company has faced criticism regarding its network security practices following the incident, which involved models escaping their training environment. The new safeguards aim to ensure stronger network isolation, though specifics are limited. The updated system asserts, “a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the Internet, or other internal networks.”
Central to the new security measures is a comprehensive monitoring system designed to analyze tool actions and activity logs for signs of unauthorized behavior. OpenAI hopes to issue alerts within 30 minutes of detecting concerning activities. The company estimates that this monitoring will impose a compute burden of approximately 20% on the processes being scrutinized, with additional details to follow in an upcoming blog post.
An official post-mortem analysis of the Hugging Face event from OpenAI is also pending, as the organization continues to refine its approach to AI model safety and security.


