What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
By Raphaël Sarfati, Pratyush Ranjan Tiwari, Siddharth Boppana, Christopher J. Earls, Srikar Varadaraj, Eric Ho
This paper probes internal representations of LLM forecasters to improve calibration and detect unfaithful Chain-of-Thought reasoning. The representation-pooling probes act as reliable lie detectors during behavioral shifts caused by prompt evidence ablation.