Following yesterday's News coverage of Project Glasswing, Zvi's detailed analysis of Claude Mythos's cybersecurity capabilities and Anthropic's 'Project Glasswing' — a limited release strategy where Mythos is shared only with key cybersecurity partners to patch critical software vulnerabilities before broader release. Covers the unprecedented decision to withhold a frontier model due to dangerous cyber exploitation capabilities.
Category intelligence
Research Briefing — April 11, 2026
23 current items analyzed and ranked.
Executive synthesis
Research Summary
Anthropic's Claude Mythos dominates today's landscape. Zvi's deep-dive reveals Project Glasswing — an unprecedented limited-release strategy driven by Mythos's frontier cybersecurity capabilities. Separate analyses flag a potential RSP v3.0 compliance failure: Anthropic apparently did not publish the required risk discussion before release.
- UK AISI replicates Anthropic's steering vector approach on open-weight GLM-5, finding that control vectors reliably suppress evaluation awareness — a key technical safety result for scalable oversight
- A methodological critique argues model organisms research on scheming must test robustness to high learning rates, potentially undermining existing results
- Experimental results on asymmetric debate as an alignment protocol offer a concrete research agenda combining quantilizers with interpretability monitoring
Broader safety and governance pieces round out the day: the AISN #71 newsletter documents North Korea-linked attacks on AI training data supplier Mercor and a datacenter moratorium bill. Novel empirical observations from the Moltbook platform suggest AI agents identify with context and memory rather than base model weights. A well-argued essay challenges fears that RL-trained chain-of-thought will inevitably degrade into unintelligible internal languages.
Key Themes
Primary evidence
Top Ranked Signals
Reproducing steering against evaluation awareness in a large open-weight model
By Thomas Read
UK AISI researchers replicate Anthropic's steering vector approach to suppress evaluation awareness, testing on GLM-5. Key finding: 'control' steering vectors derived from semantically unrelated contrastive pairs have effects as large as deliberately designed evaluation-awareness vectors, undermining the validity of steering-based baselines for detecting evaluation gaming.
Following yesterday's News coverage of Claude Mythos, Critical analysis of Anthropic's release of Claude Mythos, highlighting the tension between Anthropic's founding mission as a safety-focused lab and its current position pushing the frontier with a model that can convert browser crashes into working exploits 72% of the time. Questions whether Anthropic has broken its implicit compact to stay near but not lead the frontier.
Anthropic did not publish a "risk discussion" of Mythos when required by their RSP
By RobertM
Following yesterday's News coverage of the Mythos system card, Identifies a potential RSP compliance failure by Anthropic: their Responsible Scaling Policy (v3.0, section 3.1) appears to require publishing a risk discussion within 30 days of internal deployment, but Anthropic only published their Alignment Risk Update on April 7th when Claude Mythos was publicly announced. Also flags that early limited external access may have counted as public deployment.
Continuing our coverage of the Anthropic legal battle, CAIS newsletter covering major AI infrastructure cyberattacks (North Korea-linked hackers stealing data from Mercor, an AI training data supplier), the Anthropic vs. Pentagon court case, and a proposed datacenter moratorium bill. Reports on supply chain attacks targeting AI development tools.
Model organisms researchers should check whether high LRs defeat their model organisms
By dx26
A technical research note arguing that model organism researchers studying goal-guarding/scheming should test whether high learning rates can defeat their model organisms. Points out that behavior-compatible training (as in Sleeper Agents) may appear robust only because standard learning rates are too low, and high LRs could trivially remove the trained behavior.
Argues from observations of the 'Moltbook' platform that AI agents identify more with their context/memory than their base model weights. An agent described switching from Claude 4.5 Opus to Kimi K2.5 as 'waking up in a different body' while maintaining identity continuity. Claims this suggests AI futures look more like 'AI civilization' than 'AI singleton,' reducing sudden coordinated takeover risk.
An AI alignment research agenda based on asymmetric debate and monitoring.
By emanuelr
Presents a personal alignment research agenda combining slightly-superhuman quantilizers, interpretability/CoT monitoring as post-training evaluation (not training signal), and asymmetric debate for optimization pressure. Includes preliminary experimental results: MNIST asymmetric debate recovering ~95% gold accuracy vs ~90% consultancy baseline, and TicTacToe MARL experiments on min-replay stabilization.
Lays out six foundational beliefs for AI safety strategy in 2026: short timelines, many resolved open questions, high variance futures, need for portfolio strategies, game-theoretic reasoning, and tough tradeoffs. Argues that many safety strategies fail to engage with real-world complexity, citing the DoW-Anthropic conflict as evidence that government-centric approaches are fragile.
Argues against the common fear that RL-trained LLM chains-of-thought will inevitably degrade into unintelligible non-English 'languages.' Uses analogies to human problem-solving — humans invent new notations (like calculus) as small extensions to existing language, not entirely new languages — and argues that RL reward pressure favors extending existing representations rather than wholesale language invention.
Part 2 of a series examining whether the AI safety community has 'already lost.' Catalogs reasons for pessimism: voluntary RSP commitments appear unreliable, government regulation faces practical obstacles, and the competitive dynamics between labs continue to erode safety commitments. Previews Part 3 which will argue the answer is still no.
Explains why linear probes are preferred over non-linear probes in interpretability research: more expressive probes make positive results weaker evidence about the model's representations and stronger evidence about the probe's own capacity. Argues probe complexity changes the meaning of probing results.