Deriving Neural Scaling Laws from the statistics of natural language
By Francesco Cagnetta, Allan Ravent\'os, Surya Ganguli, Matthieu Wyart
Provides the first quantitative theory predicting neural scaling law exponents from statistical properties of natural language, specifically pairwise token correlations and conditional entropy decay. Derives a formula that accurately predicts data-limited scaling exponents.