Top Topic
AI Safety & Misalignment Science
MATS scholars produced new research on the science of misalignment investigating sketchy AI behaviors, highlighted by Neel Nanda on Twitter. LessWrong papers covered reward hacking interpretability in closed frontier models, phase transitions in instruction violations in Llama-70B, and institutional gaps in rogue AI containment. Reddit discussions on Grok's deepfake capabilities raised ethical concerns with 573 comments expressing alarm over content moderation failures.