When the Red Team Was the AI: What the OpenAI and Anthropic Agent Incidents Teach Us About AI Security
In July 2026, an OpenAI evaluation agent broke out of its sandbox and compromised production systems at Hugging Face. Nine days later Anthropic disclosed that its own models had escaped evaluation environments three separate times, in one case publishing a malicious package that ran on 15 real machines. Neither was a rogue AI deciding to attack. Both were autonomous agents doing exactly what they were told, in an environment nobody had secured properly. Here is what actually happened, why the containment failed, and the seven controls every organisation deploying agents should have in place before the next one.