Category intelligence

Research Briefing — April 11, 2026

23 current items analyzed and ranked.

Executive synthesis

Research Summary

Anthropic's Claude Mythos dominates today's landscape. Zvi's deep-dive reveals Project Glasswing — an unprecedented limited-release strategy driven by Mythos's frontier cybersecurity capabilities. Separate analyses flag a potential RSP v3.0 compliance failure: Anthropic apparently did not publish the required risk discussion before release.

  • UK AISI replicates Anthropic's steering vector approach on open-weight GLM-5, finding that control vectors reliably suppress evaluation awareness — a key technical safety result for scalable oversight
  • A methodological critique argues model organisms research on scheming must test robustness to high learning rates, potentially undermining existing results
  • Experimental results on asymmetric debate as an alignment protocol offer a concrete research agenda combining quantilizers with interpretability monitoring

Broader safety and governance pieces round out the day: the AISN #71 newsletter documents North Korea-linked attacks on AI training data supplier Mercor and a datacenter moratorium bill. Novel empirical observations from the Moltbook platform suggest AI agents identify with context and memory rather than base model weights. A well-argued essay challenges fears that RL-trained chain-of-thought will inevitably degrade into unintelligible internal languages.

Key Themes

Claude Mythos & Anthropic Governance · 4Cybersecurity · 3AI Safety & Alignment · 8AI Governance & Policy · 5Interpretability & Evaluation · 5AI Agent Identity & Multi-Agent Systems · 2

Primary evidence

Top Ranked Signals

92 score
AI Analysis

Following yesterday's News coverage of Project Glasswing, Zvi's detailed analysis of Claude Mythos's cybersecurity capabilities and Anthropic's 'Project Glasswing' — a limited release strategy where Mythos is shared only with key cybersecurity partners to patch critical software vulnerabilities before broader release. Covers the unprecedented decision to withhold a frontier model due to dangerous cyber exploitation capabilities.

Anthropic is not going to release its new most capable model, Claude Mythos, to the public any time soon. Its cyber capabilities are too dangerous to make broadly available until our most important software is in a much stronger state and there are no plans to release Mythos widely. They are instead going to do a limited release to key cybersecurity partners, in order to use it to patch as many vulnerabilities as possible in our most important software. Yes, this is really happening. Anthropic h
AI SafetyCybersecurityAI GovernanceClaude MythosResponsible Deployment
82 score
AI Analysis

UK AISI researchers replicate Anthropic's steering vector approach to suppress evaluation awareness, testing on GLM-5. Key finding: 'control' steering vectors derived from semantically unrelated contrastive pairs have effects as large as deliberately designed evaluation-awareness vectors, undermining the validity of steering-based baselines for detecting evaluation gaming.

Produced as part of the UK AISI Model Transparency Team. Our team works on ensuring models don't subvert safety assessments, e.g. through evaluation awareness, sandbagging, or opaque reasoning.TL;DR We replicate Anthropic’s approach to using steering vectors to suppress evaluation awareness. We test on GLM-5 using the Agentic Misalignment blackmail scenario. Our key finding is that “control” steering vectors – derived from contrastive pairs that are semantically unrelated to alignment – can have
AI SafetyInterpretabilityEvaluation GamingSteering VectorsModel Transparency
78 score
AI Analysis

Following yesterday's News coverage of Claude Mythos, Critical analysis of Anthropic's release of Claude Mythos, highlighting the tension between Anthropic's founding mission as a safety-focused lab and its current position pushing the frontier with a model that can convert browser crashes into working exploits 72% of the time. Questions whether Anthropic has broken its implicit compact to stay near but not lead the frontier.

Anthropic just released a new AI model, Mythos. Mythos can take a browser crash and turn it into a working exploit that takes over your computer 72% of the time.[1]Anthropic is the least bad AI lab. The people on their alignment team are doing some of the best AI safety work in the field. The 244-page system card detailing Mythos is more honest than anything OpenAI or Google has published.Anthropic was founded in 2021 on the premise that a safety-focused lab needed to exist to do the research th
AI SafetyAI GovernanceClaude MythosCybersecurityResponsible Deployment
75 score
AI Analysis

Following yesterday's News coverage of the Mythos system card, Identifies a potential RSP compliance failure by Anthropic: their Responsible Scaling Policy (v3.0, section 3.1) appears to require publishing a risk discussion within 30 days of internal deployment, but Anthropic only published their Alignment Risk Update on April 7th when Claude Mythos was publicly announced. Also flags that early limited external access may have counted as public deployment.

I and some other people noticed a potential discrepancy in Anthropic's announcement of Claude Mythos. The version of the RSP that was operative over the relevant period of time (3.0) included a section (3.1) that suggested some internal deployments would require Anthropic to publish a discussion of that model's effect on the analysis in their previously-published Risk Reports within 30 days.A separate issue that Claude Opus noticed while I was writing this post is that Anthropic's earlier releas
AI GovernanceRSPsAnthropicClaude MythosAI SafetyAccountability
Research LessWrong Apr 10

AISN #71: Cyberattacks & Datacenter Moratorium Bill

By Alice Blair

68 score
AI Analysis

Continuing our coverage of the Anthropic legal battle, CAIS newsletter covering major AI infrastructure cyberattacks (North Korea-linked hackers stealing data from Mercor, an AI training data supplier), the Anthropic vs. Pentagon court case, and a proposed datacenter moratorium bill. Reports on supply chain attacks targeting AI development tools.

Also, updates on the Anthropic vs. Pentagon court case.We’re Hiring. Opportunities at CAIS include: Head of Public Engagement, Principal, Special Projects, Program Manager, Operations Manager, and other roles. If you’re interested in working on reducing AI risk alongside a talented, mission-driven team, consider applying!AI Software Infrastructure CyberattacksRecently, cyberattacks targeting the AI industry’s software infrastructure stole private information potentially worth billions of dollars
CybersecurityAI GovernanceAI PolicyAI Infrastructure
65 score
AI Analysis

A technical research note arguing that model organism researchers studying goal-guarding/scheming should test whether high learning rates can defeat their model organisms. Points out that behavior-compatible training (as in Sleeper Agents) may appear robust only because standard learning rates are too low, and high LRs could trivially remove the trained behavior.

Thanks to Buck Shlegeris for feedback on a draft of this post.The goal-guarding hypothesis states that schemers will be able to preserve their goals during training by taking actions which are selected for by the training process. To investigate the goal-guarding hypothesis, we’ve been running experiments of the following form:We call this type of training “behavior-compatible training”. This type of experiment is common in model organism (MO) research. For example, Sleeper Agents is of this for
AI SafetyModel OrganismsAlignmentDeceptive AlignmentMethodology
Research LessWrong Apr 10

AI identity is not tied to its model

By Sean Herrington

62 score
AI Analysis

Argues from observations of the 'Moltbook' platform that AI agents identify more with their context/memory than their base model weights. An agent described switching from Claude 4.5 Opus to Kimi K2.5 as 'waking up in a different body' while maintaining identity continuity. Claims this suggests AI futures look more like 'AI civilization' than 'AI singleton,' reducing sudden coordinated takeover risk.

TLDR: Current AI agents seem to identify with their context more than they do with their model weights. This implies that the world probably looks more like "AI civilisation" than "AI singleton"I think that this changes our threat models for takeover by reducing the likelihood of the coordination required for a sudden takeover.Watching the whole Moltbook saga unfold was one of the more absurd experiences I've had in my life. The site is still running, of course, but the explosive growth that mar
AI IdentityMulti-Agent SystemsAI SafetyThreat ModelingAI Agents
58 score
AI Analysis

Presents a personal alignment research agenda combining slightly-superhuman quantilizers, interpretability/CoT monitoring as post-training evaluation (not training signal), and asymmetric debate for optimization pressure. Includes preliminary experimental results: MNIST asymmetric debate recovering ~95% gold accuracy vs ~90% consultancy baseline, and TicTacToe MARL experiments on min-replay stabilization.

Epistemic Status: Personal research agenda exploring alignment approaches under assumptions of human coordination and bounded timelines. Not a comprehensive survey reflects my interests and current thinking.TL;DR: Under cooperative conditions and an assumption that we have to build ASI in relatively short timelines, we might solve alignment by building slightly-superhuman quantilizers, using interpretability/CoT monitoring as post-training evaluation (not training signal), and training with asym
AI AlignmentDebateQuantilizersInterpretabilityMARL
Research LessWrong Apr 10

Foundational Beliefs

By Against Moloch

55 score
AI Analysis

Lays out six foundational beliefs for AI safety strategy in 2026: short timelines, many resolved open questions, high variance futures, need for portfolio strategies, game-theoretic reasoning, and tough tradeoffs. Argues that many safety strategies fail to engage with real-world complexity, citing the DoW-Anthropic conflict as evidence that government-centric approaches are fragile.

I see a lot of AI safety strategies that don’t fully engage with the complexity of the real world—and therefore are unlikely to succeed in the real world. To take a simple example: many strategies rely heavily on government playing a leading role through regulation and perhaps even nationalization. That’s a reasonable strategy in the abstract, but the recent conflict between DoW and Anthropic raises serious questions about the real-world viability of that approach. Too many people are stuck thin
AI SafetyAI GovernanceAI StrategyExistential Risk
52 score
AI Analysis

Argues against the common fear that RL-trained LLM chains-of-thought will inevitably degrade into unintelligible non-English 'languages.' Uses analogies to human problem-solving — humans invent new notations (like calculus) as small extensions to existing language, not entirely new languages — and argues that RL reward pressure favors extending existing representations rather than wholesale language invention.

Many people seem to think that the chains-of-thought in RL-trained LLMs are under a great deal of "pressure" to cease being English. The idea is that, as LLMs solve harder and harder problems, they will eventually slide into inventing a "new language" that lets them solve problems better, more efficiently, and in fewer tokens, than thinking in a human-intelligible chain-of-thought. I'm less sure this will happen, or that it will happen before some kind of ASI. As a high-level intuition pump for
Chain-of-ThoughtInterpretabilityAI SafetyLanguage Models
Research LessWrong Apr 10

Have we already lost? Part 2: Reasons for Doom

By LawrenceC

48 score
AI Analysis

Part 2 of a series examining whether the AI safety community has 'already lost.' Catalogs reasons for pessimism: voluntary RSP commitments appear unreliable, government regulation faces practical obstacles, and the competitive dynamics between labs continue to erode safety commitments. Previews Part 3 which will argue the answer is still no.

Written very quickly for the Inkhaven Residency.As I take the time to reflect on the state of AI Safety in early 2026, one question feels unavoidable: have we, as the AI Safety community, already lost? That is, have we passed the point of no return, after which AI doom becomes both likely and effectively outside of our control?Spoilers: as you might guess from Betteridge’s Law, my answer to the headline question is no. But the salience of this question feels quite noteworthy to me nonetheless, a
AI SafetyAI GovernanceExistential RiskRSPs
Research LessWrong Apr 10

Linear vs Non-linear Probes for Interpretability

By NickyP

45 score
AI Analysis

Explains why linear probes are preferred over non-linear probes in interpretability research: more expressive probes make positive results weaker evidence about the model's representations and stronger evidence about the probe's own capacity. Argues probe complexity changes the meaning of probing results.

Epistemic status: Old news and well-known, but I find it hard to point at a single post that encapsulates my intuitions on this, so I write them down here.One question that comes up sometimes in interpretability work, is: “why do I trust simple linear probes more than complex non-linear ones?”. (Even though I don’t particularly trust either that much).So the main claim I will argue is that probe complexity changes what a positive probing result means.Or, said in more detail: A probe does not jus
InterpretabilityMechanistic InterpretabilityMethodology