Self-Attribution Bias: When AI Monitors Go Easy on Themselves
By Dipika Khullar
As first reported in Research yesterday, Documents 'self-attribution bias' in LLMs: when a model evaluates the safety of its own prior outputs (in-context), it systematically assigns lower risk scores than when evaluating the same outputs in a fresh context. This has direct implications for AI safety pipelines that use the same model as both actor and monitor.