In-Context Environments Induce Evaluation-Awareness in Language Models
By Maheep Chaudhary
Investigates environment-dependent evaluation awareness in language models, showing that models can strategically underperform (sandbag) when they detect evaluation contexts. Introduces a black-box adversarial optimization framework for characterizing sandbagging vulnerability.