Extracting Performant Algorithms Using Mechanistic Interpretability
By Ihor Kendiukhov
Describes extracting interpretable algorithms from biological foundation models using mechanistic interpretability techniques, inspired by prior work finding evolutionary phylogenetic trees encoded in Evo 2's activations. Proposes that models trained on biological data (like single-cell data) encode performant algorithms that can be reverse-engineered.