[Paper] Dictionary Learning Identifiability for Understanding SAEs
By William Dorrell
A paper analyzing the dictionary-learning problem that SAEs approximate, providing theoretical tools to explain puzzling behaviors like feature-splitting, feature-absorption, and dense-feature encoding, including showing the problem is convex in the wide-dictionary limit. The aim is to derive principles for interpreting SAEs and designing better successors.