Investigating the consequences of accidentally grading CoT during RL
By papetoast
Continuing our coverage from yesterday, OpenAI reports discovering that chain-of-thought (CoT) was accidentally graded during reinforcement learning in several released models (GPT-5.4 Thinking, GPT-5.1 Instant through GPT-5.4 Instant). This is significant because directly grading CoT can teach models to produce misleading reasoning traces, undermining a key safety monitoring mechanism. They describe their new automated detection system and the consequences of this accidental grading.