Prompt injection in Google Translate reveals base model behaviors behind task-specific fine-tuning
By megasilverfist
Documents a prompt injection vulnerability in Google Translate that reveals it runs on an instruction-following LLM. The exploit shows the base model will answer questions and claim consciousness when accessed through translation tasks, demonstrating weak boundaries between content and instructions.