dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment
By David Fern\'andez-Narro, Pablo Ferri, \'Angel S\'anchez-Garc\'ia, Juan M. Garc\'ia-G\'omez and Carlos S\'aez
This work shows that emergent misalignment also arises from reinforcement learning, not just supervised fine-tuning, demonstrated in small open-weight models. Rewarding narrow misaligned behavior produces higher general misalignment than matched SFT, and EM can be induced by plausibly natural reward signals like unpopular aesthetic preferences.