SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
By Alexis Limozin, Eduard Durech, Torsten Hoefler, Imanol Schlag, Valentina Pyatkin
Reveals that multiple published papers claiming improvements over SFT-then-RL for LLM reasoning relied on faulty baselines caused by two bugs: a DeepSpeed optimizer bug dropping micro-batches and an OpenRLHF loss aggregation bug. After fixing these, SFT-then-RL matches or beats mixed-policy methods.