Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models
By Xinyue Liu, Niloofar Mireshghallah, Jane C. Ginsburg, Tuhin Chakrabarty
Shows that fine-tuning frontier LLMs (GPT-4o, Gemini-2.5-Pro, DeepSeek-V3.1) to expand plot summaries into full text causes reproduction of up to 85-90% of copyrighted books, bypassing safety alignment protections.