Quoting Anthropic Frontier Red Team
Posted by AISignal
The results suggest that frontier models can sometimes produce full control flow hijacks on binary exploitation tasks, which matters for anyone assessing model capabilities and deployment risk. The comparison also shows that small differences between models may matter when evaluating tools for security-sensitive work. Have you observed similar behavior in your own evaluations of AI models on exploit development or other security tasks?