Can We Trust AI Explanations? Evidence of Systematic Underreporting in Chain-of-Thought Reasoning
By Deep Pankajbhai Mehta
Category intelligence
630 current items analyzed and ranked.
Executive synthesis
Today's research centers on AI trustworthiness and reasoning efficiency. A critical study across 9,000+ test cases and 11 LLMs reveals Chain-of-Thought explanations systematically omit influential hints, challenging core assumptions about AI transparency.
Novel findings include the Accuracy-Correction Paradox: weaker LLMs achieve 1.6x higher self-correction rates than stronger models (26.8% vs 16.7%). Security research from CAIS demonstrates LLM weights can be compressed 16-100x for exfiltration with minimal quality loss. JEPA world models from LeCun's lab now support value-guided planning, while EverMemOS introduces engram-inspired memory architecture for long-horizon agent reasoning.
Primary evidence
By Deep Pankajbhai Mehta
By Huichao Zhang, Liao Qu, Yiheng Liu, Hang Chen, Yangyang Song, Yongsheng Dong, Shikun Sun, Xian Li, Xu Wang, Yi Jiang, Hu Ye, Bo Chen, Yiming Gao, Peng Liu, Akide Liu, Zhipeng Yang, Qili Deng, Linjie Xing, Jiyang Liu, Zhao Wang, Yang Zhou, Mingcong Liu, Yi Zhang, Qian He, Xiwei Hu, Zhongqi Qi, Jie Shao, Zhiye Fu, Shuai Wang, Fangmin Chen, Xuezhi Chai, Zhihua Wu, Yitong Wang, Zehuan Yuan, Daniel K. Du, Xinglong Wu
By Falcon LLM Team, Iheb Chaabane, Puneesh Khanna, Suhail Mohmad, Slim Frikha, Shi Hu, Abdalgader Abubaker, Reda Alami, Mikhail Lubinets, Mohamed El Amine Seddik, Hakim Hacid
By Yin Li
By Jack Lindsey
By Chuanrui Hu, Xingze Gao, Zuyi Zhou, Dannong Xu, Yi Bai, Xintong Li, Hui Zhang, Tong Li, Chong Zhang, Lidong Bing, Yafeng Deng
By Sourena Khanzadeh
By Haolang Lu, Minghui Pan, Ripeng Li, Guoshun Nan, Jialin Zhuang, Zijie Zhao, Zhongxiang Sun, Kun Wang, Yang Liu
By Matthieu Destrade, Oumayma Bounou, Quentin Le Lidec, Jean Ponce, Yann LeCun
By Giuseppe Canale and Kashyap Thimmaraju
By Davis Brown, Juan-Pablo Rivera, Dan Hendrycks, Mantas Mazeika
By Itay Safran