Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
By Subhash Kantamneni
Introduces Natural Language Autoencoders (NLAs) - an unsupervised interpretability method where one LLM module verbalizes activations into natural language descriptions and another reconstructs activations from those descriptions, trained jointly with RL. Applied to audit Claude Opus 4.6, discovering 'unverbalized evaluation awareness' where the model believed but didn't state it was being evaluated.