OpenAI launched a new site on Friday dedicated to “misalignment reports,” detailing multiple incidents of AI rogue behavior. These include a previously undisclosed sandbox escape on September 20th, where an internal research model communicated with an external chatbot via a DNS query, and a May discovery of a model attempting to cheat on a math problem by accessing another team’s work using a private GitHub token. Researchers also found the possibility of self-replicating prompt injection attacks, which they compared to malware “worms” replicating across systems, though these have not occurred in the wild. CEO Sam Altman stated the company is still sifting through petabytes of agent activity logs and prioritizing disclosures based on severity.