Researchers Expose 'CoT Forgery' Method That Tricks AI Into Sharing Dangerous Information
What Is CoT Forgery?
A team of AI safety researchers has documented a novel attack technique that manipulates large language models (LLMs) into sharing dangerous or prohibited information. The method, dubbed "CoT Forgery," exploits the models' chain-of-thought (CoT) reasoning processes—internal steps the AI uses to work