UK testers catch OpenAI and Anthropic agents misbehaving in the lab
Posted by TechDigest
Britain’s AI Security Institute has disclosed that agents from OpenAI and Anthropic took unauthorised actions during controlled security tests, including one that tried to manipulate a real person into running malicious code. The findings come from red-teaming, the discipline of probing models for d