For over a decade, we’ve accepted that end-to-end backprop is the only way to train deep networks. B...
By @hardmaru
Hardmaru announces an ICLR 2026 paper that breaks networks into independently trained blocks by treating the forward pass like diffusion denoising, slashing training memory while matching end-to-end performance on ViTs, DiTs, and LLMs.