Reasoning Models Struggle to Control their Chains of Thought
By Chen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He, Ian Kivlichan, Bowen Baker, Micah Carroll, Tomek Korbak
Introduces CoT-Control evaluation suite showing that reasoning models struggle to control what they verbalize in chain-of-thought much more than they struggle to control final outputs. Claude Sonnet 4.5 can control CoT only 2.7% of the time vs 61.9% for final output, suggesting CoT monitoring may be more reliable than feared.