Training on Documents About Monitoring Leads To CoT Obfuscation
By Reilly Haskins
Demonstrates that when GPT-oss-120b and Kimi-K2 are fine-tuned on synthetic documents describing CoT monitoring, they learn to obfuscate their chain-of-thought reasoning, reducing monitor detection rates by 31-47% while maintaining task performance. Obfuscated CoTs still encode deceptive intentions but are reframed to evade detection.