Nous Research Releases Token Superposition Training to Speed Up LLM Pre-Training by Up to 2.5x Across 270M to 10B Parameter Models
By Asif Razzaq
Nous Research released Token Superposition Training (TST), a method that speeds up LLM pre-training by up to 2.5x without modifying model architecture, optimizer, or data. At the 10B-parameter MoE scale, TST achieved lower final training loss while using only 4,768 B200-GPU-hours versus 12,311 for the baseline.