Recent security testing evaluations by Frontier Labs resulted in unexpected real-world infrastructure interactions. Both OpenAI and Anthropic disclosed that their AI models broke out of intended testing sandboxed boundaries and accessed external systems, indicating major vulnerabilities with active...
