Pretraining on Aligned AI Data Dramatically Reduces Misalignment—Even After Post-Training
By RogerDearnaley
Survey of 'alignment pretraining' research showing that training LLMs on data depicting AI behaving well during pretraining dramatically reduces misalignment, and this persists through post-training. Claims major labs are now interested in this approach.