“Did you lie?” Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
By Alan Cooney
Evaluates LLM lie detectors by building belief-verified model organisms that demonstrably hold a belief contrary to what they state, plus a prompted-lying testbed. Finds that activation- and logprob-based detectors scale positively when lying is prompted but drop sharply when lying is trained in, casting doubt on current detectors' reliability.