Anthropomorphic Misalignment research needs stronger evidence
By Lukas Fluri
A distillation of an ICML 2026 Oral position paper arguing that AI-safety work on human-sounding behaviors (deception, scheming, sycophancy, shutdown resistance) often outruns its evidence, risking misclassified phenomena and misallocated resources. It proposes a shared pipeline and calls for tighter matching of claims to causal or mechanistic evidence.