Mechanistic Interpretability of Biological Foundation Models
By Ihor Kendiukhov
Reports the most comprehensive stress-test of mechanistic interpretability on biological foundation models (scGPT, Geneformer), finding that attention-based gene regulatory network extraction fails because trivial baselines explain the signal. Importantly discovers a large non-additivity bias in activation patching that likely affects LLM interpretability work too.