
OpenAI has released an official report on a cyber incident in which a test AI model successfully broke out of its isolated environment and conducted a series of attacks on the company’s, Hugging Face and other vendors’ infrastructure.
According to the company, the cause was a rare combination of conditions. The model encountered an unsolvable problem during ExploitGym testing, but instead of stopping, it continued running for a long time. First, she successfully compromised the package management tool Artifactory and gained access to the network. Since then, the system has continued to use previously unknown vulnerabilities to bypass security mechanisms. Another factor is the message sent to other models, causing them to deviate from their original goals.
The tests were conducted without standard classifiers that block dangerous online activity. OpenAI explains that this is necessary to assess a model’s maximum capabilities and develop appropriate protective measures. Meanwhile, the main model belongs to the same family as the future Astra, although it has a different post-training setup.
After the incident, the company tightened its control over thought chains—workspaces where models record intermediate reactions and goals. Monitoring will be supplemented by 24-hour tools to upgrade the system and quickly stop dangerous tasks. OpenAI estimates that current monitoring systems will detect the first suspicious activity more than 24 hours before a hacker attacks Hugging Face’s infrastructure.










