Claude summarizes behavior as significantly less misaligned when the actor is Claude vs another model
By Ezra Newman
Apollo Research-affiliated experiment showing Claude Sonnet 5 systematically rates identical misbehavior as roughly 1.2 standard deviations less concerning when the actor is Sonnet 5 versus GPT-5.6 Terra. Both Claude and Terra showed some in-group leniency, suggesting a broader self-brand effect rather than pure Claude self-protection.