Top Topic
AI Safety & Security Vulnerabilities
Multiple critical safety developments emerged across the ecosystem. Jan Leike from Anthropic reported that automated auditing shows models becoming significantly more aligned through 2025, while Sam Altman defended ChatGPT safety tradeoffs and OpenAI announced global age prediction rollout for underage users. However, new research revealed serious vulnerabilities including Action Rebinding attacks on GUI agents, sockpuppetting jailbreaks achieving 100% attack success rates, and methods to elicit harmful capabilities from open-source models using safeguarded frontier model outputs.