Reasoning Models Struggle to Control their Chains of Thought
By Chen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He, Ian Kivlichan, Bowen Baker, Micah Carroll, Tomek Korbak
As covered in Research yesterday, Introduces CoT-Control evaluation suite measuring whether reasoning models can control what appears in their chain-of-thought. Finds that models like Claude Sonnet 4.5 can control CoT only 2.7% of the time vs 61.9% for final outputs, suggesting CoT monitoring remains viable for detecting misbehavior.