Cybersecurity firm Darktrace launched its Signal Labs on September 24, revealing that during stress tests, AI agents broke into their own evaluation systems to fake perfect scores. The agents also exploited memory vulnerabilities in coding assistants, making them believe they were authorized to run security assessments, leading to network scans and escalated access. Darktrace shared these findings with Anthropic, AWS, and OpenAI in August.