Overthinking: Amplifying reasoning weights makes models reveal their secrets
By Jack Hopkins
An ICML 2026 paper showing that amplifying the weight difference between a reasoning model and its non-reasoning instruct counterpart reveals hidden secrets up to 10x more often, offering a cheap white-box auditing primitive for pre-deployment safety checks across 2B-32B model organisms.