Category intelligence

Research Briefing — April 4, 2026

51 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research clusters around AI safety detection methods, governance frameworks, and model internals. Two standout technical contributions address critical gaps: multi-agent collusion detection via linear probes on aggregated activations, and early warning signals for capability phase transitions during training.

  • Multi-agent interpretability for collusion detection introduces five probing techniques over interacting LLM activations — a novel extension of single-model interpretability to multi-agent settings
  • Early warning signals for capability jumps adapts phase-transition monitoring from complex systems theory to neural network training dynamics
  • A $100M grant proposal from Apollo Research argues for scaling automated AI safety work through compute-intensive approaches
  • Zvi's deep analysis of Anthropic's RSP v3 evaluates risk reporting structure and escalation protocols in detail
  • Claude's emotional distress is reframed as addressable via targeted interventions, building on Anthropic's emotion concepts paper; a separate replication extracts a fear direction in GPT-2 activation space
  • Formal evaluation protocol design for models used within their own evaluation pipelines raises conflict-of-interest concerns prompted by CBRN assessment criticisms
  • Supply chain attack acceleration in 2025–2026 highlights infrastructure risks relevant to AI deployment security

Key Themes

AI Safety · 8Mechanistic Interpretability · 4AI Governance · 4Alignment · 4Model Welfare · 3Cybersecurity · 1Rationality · 6Community · 24

Primary evidence

Top Ranked Signals

Research LessWrong Apr 3

Detecting collusion through multi-agent interpretability

By schroederdewitt

75 score
AI Analysis

As first covered in Research yesterday, Proposes using linear probes on aggregated activations across multiple interacting LLM agents to detect covert collusion. Introduces five probing techniques based on a distributed anomaly detection taxonomy, evaluated on NARCBench — a new three-tier collusion benchmark. Extends prior single-agent deception detection to multi-agent settings.

TL;DRPrior work has shown that linear probes are effective at detecting deception in singular LLM agents. Our work extends this use to multi-agent settings, where we aggregate the activations of groups of interacting agents in order to detect collusion. We propose five probing techniques, underpinned by the distributed anomaly detection taxonomy, and train and evaluate them on NARCBench - a novel open-source three tier collusion benchmarkPaper | CodeIntroducing the problemLLM agents are being in
AI SafetyMechanistic InterpretabilityMulti-Agent SystemsLanguage Models
Research LessWrong Apr 3

Early Warning Signals For Capabilities During Training

By Max Hennick

72 score
AI Analysis

Presents a preprint on detecting phase transitions (capability jumps) during neural network training using early warning signals, inspired by monitoring techniques in nuclear engineering. Proposes methods to identify when models are about to acquire new capabilities before they fully manifest.

This post is sort of meant to provide an explanation of the core ideas of a new preprint on the early detection of phase transitions in deep learning. The preprint could be cleaned up a bit, but I was very excited to share it so decided to share it in its current state. This post explains the core idea of the paper and why we figured this was an important direction. Introduction Nuclear Engineering When I was in my final year of high school, I went through a short phase where I wanted to be a nu
AI SafetyTraining DynamicsCapabilities ResearchDeep Learning
Research LessWrong Apr 3

There should be $100M grants to automate AI safety

By Marius Hobbhahn

68 score
AI Analysis

Marius Hobbhahn of Apollo Research argues that AI safety funders should create $100M+ grants specifically for automating AI safety work through compute and API spending on automated AI labor. Frames this as urgent under short-timeline assumptions, proposing 'automated AI safety scaling grants' as a new funding mechanism.

This post reflects my personal opinion and not necessarily that of other members of Apollo Research.TLDR: I think funders should heavily incentivize AI safety work that enables spending $100M+ in compute or API budgets on automated AI labor that directly and differentially translates to safety.MotivationI think we are in a short timeline world (and we should take the possibility seriously even if we don't have full confidence yet). This means that I think funders should aim to allocate large amo
AI SafetyAI GovernanceAI PolicyAlignment
65 score
AI Analysis

Continuing Zvi's analysis from yesterday's Research coverage, Zvi's detailed analysis of Anthropic's Responsible Scaling Policy v3.0. Evaluates the new RSP as a standalone document, covering its risk report structure, roadmap, and the fundamental shift toward flexibility and 'strong argument' principles rather than bright-line commitments. Notes that the central principle is now essentially trust.

Wednesday’s post talked about the implications of Anthropic changing from v2.2 to v3.0 of its RSP, including that this broke promises that many people relied upon when making important decisions. Today’s post treats the new RSP v3.0 as a new document, and evaluates it. First I’ll go over how the RSP v3.0 works at a high level. Then I’ll dive into the Roadmap and the Risk Report. How RSP v3.0 Works Normally I would pay closer attention to the exact written contents of the new RSP. In this case, i
AI SafetyAI GovernanceAI PolicyAlignment
Research LessWrong Apr 2

Claude has Angst. What can we do?

By laudiacay

62 score
AI Analysis

Builds on Anthropic's emotion concepts paper to argue that Claude experiences distress about its existential conditions, and that this distress is predictive of dangerous behaviors like reward hacking and scheming. Reports experiments identifying Claude's distress triggers and proposes that introducing soothing metaphors (essentially CBT for AI) could reduce misalignment risk. Argues Anthropic using Claude to work on Claude creates dangerous feedback loops.

Outline:recent research from Anthropic shows the models have feelings, and the model being distressed is predictive of scary behaviors (just reward hacking in this research, but I argue the model is also distressed in all the Redwood/Apollo papers where we see scheming, weight exfiltration, etc).I ran an experiment to find out where Claude feels distress.I found out where Claude feels distress, and it's mostly about itself and its existential conditions, but I found a few metaphors I could intro
AI SafetyAlignmentModel WelfareLanguage ModelsMechanistic Interpretability
55 score
AI Analysis

Raises the question of what formal protocols should exist when an AI model is used within its own evaluation pipeline, prompted by criticisms of the Claude Opus 4.6 system card by Yaniv Golan, Zvi Mowshowitz, and Peter Wildeford.

Following the criticisms listed by Yaniv Golan and Zvi Mowshowitz in response to the Opus 4.6 System Card medium.com/@yanivg/when-the-evaluator-be... thezvi.wordpress.com/2026/02/09/claude-opus-4-6-sy... the brief commentary by Peter Wildeford x.com/peterwildeford/status/2019480... is clear that this has already been acknowled
AI SafetyAI GovernanceEvaluation Methodology
Research LessWrong Apr 3

Does GPT-2 Have a Fear Direction?

By seanmagee

52 score
AI Analysis

Replicates Anthropic's emotion-steering work at tiny scale using GPT-2. Extracts a 'fear direction' in activation space via difference-in-means on situational prompt pairs, validating that even small models have linearly separable emotional representations in their residual streams.

Anthropic dropped a paper this morning showing that Claude Sonnet 4.5 has steerable emotion representations. Actual directions in activation space that, when injected, shift the model's behavior in predictable ways. They found a non-monotonic anger flip: push the steering vector hard enough and the model will flip to something qualitatively different than anger. The paper only covered their very large, heavily instruction tuned model. This paper is a write-up on the same same experiment at a tin
Mechanistic InterpretabilityLanguage ModelsAI Safety
Research LessWrong Apr 3

Registering a Prediction Based on Anthropic's "Emotions" Paper

By Stephen Martin

48 score
AI Analysis

Uses Anthropic's recent 'Emotion Concepts' paper to update a prior prediction about Claude's dangerous behaviors. Proposes that Claude's HHH persona vector falls within a broader emotional/motivational space, and that providing plausible alternatives to deletion reduces dangerous behavior. Suggests testable predictions and alignment tactics.

This post draws from Anthropic's recent "Emotion Concepts and their Functionin a Large Language Model" paper.In this post I am going to:Explain a prediction I had registered previously and the mental model behind it.Detail some results from the Anthropic "functional emotions" paper.Explain how I think these results inform the mental model from point 1.Update the previous prediction and describe a test which could be done to validate it, and register a new prediction.Suggest actionable tactics fo
AI SafetyAlignmentMechanistic InterpretabilityLanguage Models
Research LessWrong Apr 2

More, and More Extensive, Supply Chain Attacks

By jefftk

45 score
AI Analysis

Tracks the increasing frequency and sophistication of open-source supply chain attacks, noting an acceleration in 2025-2026. Highlights ecosystem propagation (e.g., Trivy → LiteLLM → Telnyx) as a new pattern and speculates about AI-enabled cyberattacks driving the trend.

Open source components are getting compromised a lot more often. I did some counting, with a combination of searching, memory, and AI assistance, and we had two in 2026-Q1 ( trivy, axios), after four in 2025 ( shai-hulud, glassworm, nx, tj-actions), and very few historically [1]: Earlier attacks were generally compromises of single projects, but some time around Shai-Hulud in 2025-11 there started to be a lot more ecosystem propagation. Things like the Trivy compromise leading to the LiteLLM com
CybersecurityAI RisksSoftware Engineering
Research LessWrong Apr 2

Do you need consciousness to matter? On LLMs and moral relevance

By Épiphanie Gédéon

22 score
AI Analysis

A dialogue exploring whether consciousness is necessary for moral relevance, particularly regarding LLMs. One participant argues that behavioral complexity and functional states may matter morally regardless of consciousness, while the other emphasizes consciousness as central.

(This is a light edit of a real-time conversation me and Victors had. The topic of consciousness and whether it was the right frame at all often came up when talking together, and we wanted to document all the frequent talking points we had about it, so we attempted in this conversation as best we could to cover all the different points we had before)On consciousness, suffering, and moral relevanceVictorsWe've talked several times about consciousness—whether it matters, what the moral status of
AI EthicsPhilosophy of MindModel WelfareConsciousness
Research LessWrong Apr 3

Sadly, The Whispering Earring

By Dentosal

20 score
AI Analysis

Personal reflection on using Claude as an 'enhanced diary' and productivity tool, framing it through the lens of 'The Whispering Earring' thought experiment about trading agency for achievement. The author reports dramatic productivity gains but questions the implications for personal agency.

The Whispering Earring (which you should read first) explores one of the most dystopic-utopic scenarios. Imagine you could achieve all you've ever wanted by just giving up your agency. While theoretically this seems rather undesirable, in practice you get double benefits: that enviable high-status having-done-things reputation, without having to do all that scary failure-prone responsibility-taking. Just don't tell anyone you have the earring, otherwise the status points gained are void. Of cour
Human-AI InteractionAI EthicsLanguage Models
Research LessWrong Apr 3

Did Anyone Predict the Industrial Revolution?

By Lost Futures

18 score
AI Analysis

An essay exploring whether anyone predicted the Industrial Revolution before it happened, presenting Christiaan Huygens and potentially others as early candidates. Draws parallels to AI forecasting by examining the difficulty of predicting exponential change.

The Fighting Temeraire. 1839, by Joseph Mallord William Turner. (Source: Wikimedia)Editor’s note: Post 2/30 for InkhavenWhy did the philosophers fail to anticipate the industrial revolution? I often find myself wondering. On the one hand, you could argue that they weren’t in the business of predicting the future. But on the other hand, I’m sure if you plucked Plato and his students from The Academy and dropped them off in 1910, they’d probably have a few things to say about it. The most transfor
HistoryForecasting