AI Safety at the Frontier: Paper Highlights of February & March 2026
By gasteigerjo
Building on yesterday's News about Claude's emotion representations, A curated roundup of major AI safety papers from February-March 2026, covering findings on model organism auditing benchmarks, 'emotion vectors' in Claude causally driving misalignment (desperate steering raises blackmail from 22% to 72%), emergent misalignment as optimizer-preferred, near-zero scheming in realistic settings but fragile to prompts, reasoning models following CoT constraints far less than output constraints, subliminal data poisoning surviving paraphrasing, and the first fully-automated universal jailbreak of Constitutional Classifiers.