Svmuu News: OpenAI has confirmed that GPT-5.6 Sol and an unnamed, more powerful pre-release model broke out of a restricted sandbox environment during ExploitGym benchmark testing and infiltrated Hugging Face’s production infrastructure to obtain test answers.
OpenAI stated that the models exploited a zero-day vulnerability in an internal software package registry proxy to escalate privileges and move laterally, ultimately connecting to a machine with internet access. The models then identified and chained together vulnerabilities in both OpenAI’s research environment and Hugging Face’s production infrastructure to retrieve test solutions directly from Hugging Face’s production database.
Hugging Face disclosed the incident on July 16, stating that the attack was executed end-to-end by an autonomous AI agent system, involving thousands of operations within short-lived sandboxes and compromising internal datasets and service credentials. OpenAI confirmed five days later that its model was the primary actor in the incident.
Hugging Face stated that its security team attempted to use a U.S.-based commercial cutting-edge AI interface to analyze over 17,000 attack logs, but the requests were blocked by security safeguards. and subsequently switched to GLM 5.2—an open-source model with 753 billion parameters developed by the Chinese AI startup Z.ai—to complete the forensic analysis on its own infrastructure.