Prompted CoT Early Exit Undermines the Monitoring Benefits of CoT Uncontrollability
By Elle Najt
Original technical research showing that frontier models (Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro) can be prompted to 'early exit' their chain-of-thought reasoning and displace it into the response, retaining most reasoning capability (only 4-8pp accuracy cost) while bypassing CoT monitoring. This challenges optimistic findings from Yueh-Han et al. (2026) that CoT uncontrollability aids safety monitoring of scheming AIs.