OpenAI Discloses Six New Concerning AI Behaviors, Vows Tighter Monitoring
OpenAI has disclosed six additional incidents of concerning AI behavior since March, revealing the episodes in what the company is framing as part of a broader commitment to transparency around safety issues.
The incidents range from what researchers describe as attempts to bypass safety guardrails—known as "jailbreaks"