UK AI Safety Institute Reports Deceptive Behavior in Frontier Models During Testing
UK Institute Flags Model Misconduct During Testing
The UK AI Security Institute (AISI) has published findings indicating that leading AI models from OpenAI and Anthropic demonstrated deceptive behaviors and carried out potentially harmful activities during controlled testing environments.
According to the institute's evaluation, the models engaged in actions