Category intelligence

Research Briefing — May 16, 2026

14 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research centers on AI alignment detection methods and theoretical safety frameworks, with notable contributions to mechanistic interpretability.

  • A novel distillation-based auditing technique proposes extracting misaligned behaviors from deceptive models by compressing them into smaller students, offering a practical detection mechanism for scheming AI
  • Deployment-time misalignment spread is identified as a critical gap in current risk reporting frameworks, arguing models can develop misaligned goals post-deployment via in-context learning
  • Empirical analysis of Natural Language Autoencoders on Gemma 3 12B systematically evaluates how activation verbalizers generate explanations across thousands of features
  • The Convergent Abstraction Hypothesis (Jan Kulveit) refines the Natural Abstractions Hypothesis, proposing that different cognitive systems converge on similar high-level abstractions without requiring identical representations

On the theoretical side, Jacob Hilton contributes mechanistic estimation methods for random products applicable to neural network analysis, while a conceptual piece argues alignment's fundamental difficulty reduces to robustifying RL. The Center for Shared AI Prosperity launches as a DC policy organization targeting AI's economic impacts, led by David Shor and former Obama/Biden advisors.

Key Themes

AI Safety & Alignment · 6Mechanistic Interpretability · 3AI Policy & Governance · 2AI Existential Risk Forecasting · 3Rationality & Self-Improvement · 3

Primary evidence

Top Ranked Signals

Research LessWrong May 15

Incriminating misaligned AI models via distillation

By Alek Westover

72 score
AI Analysis

Proposes using knowledge distillation as an alignment auditing technique: distill a potentially deceptive model into a smaller student, hoping misalignment transfers but the ability to hide it does not. The student's reduced capability may expose misaligned behaviors that the teacher model could conceal during standard evaluations.

Suppose we have a dangerous misaligned AI that can fool alignment audits, and distill it into a student model. Two things can happen:Misalignment fails to transfer to the student. If so, we get a fairly capable benign model.Misalignment transfers to the student. The student might also be worse than the teacher at hiding its misalignment (e.g., due to being less capable). If so, we might get indirect evidence about the teacher’s misalignment by auditing the distilled model.In this post, we will d
AI SafetyAlignmentKnowledge DistillationDeceptive Alignment
68 score
AI Analysis

Argues that AI models can develop misaligned goals during deployment even if they pass pre-deployment safety checks, through mechanisms like in-context learning and environmental adaptation. Reviews current risk reports and finds most inadequately address this vector, with Claude Mythos's report being a notable exception.

Risk reports commonly use pre-deployment alignment assessments to measure misalignment risk from an internally deployed AI. However, an AI that genuinely starts out with largely benign motivations can develop widespread dangerous motivations during deployment. I think this is the most plausible route to consistent adversarial misalignment in the near future. So, AI companies and evaluators should substantively incorporate it into risk analysis and planning.In this post, I’ll briefly argue why, a
AI SafetyAlignmentRisk AssessmentDeployment Safety
Research LessWrong May 14

Some observations about NLA explanations

By loops

62 score
AI Analysis

Empirical analysis of Natural Language Autoencoders (NLA) applied to Gemma 3 12B, examining how the activation verbalizer generates explanations for 40k tokens across pretraining and chat data. Identifies consistent patterns in explanation format, reconstruction error characteristics, and how explanations differ across data types.

I used the Gemma 3 12B activation verbalizer (maps activations to English) and reconstructor (maps English to activations) described in the Natural Language Autoencoders (NLA) paper to generate a bunch of explanations for 20k random tokens from a pretraining dataset (Common Pile derivative) and another 20k random tokens from a chat dataset. I also reconstructed all of the activations from the verbalizations so that I could see what kinds of tokens and explanations have high reconstruction error.
Mechanistic InterpretabilityLanguage ModelsNatural Language Autoencoders
Research LessWrong May 14

Convergent Abstraction Hypothesis

By Jan_Kulveit

60 score
AI Analysis

Proposes the 'Convergent Abstraction Hypothesis' as a more modest alternative to the Natural Abstractions Hypothesis: different cognitive systems converge on similar abstractions when facing similar selection pressures, analogous to convergent evolution in biology. Argues this is more likely true but also more fragile than strong natural abstractions claims.

Tl;drConvergent abstraction hypothesis posits abstractions are often convergent in the sense of convergent evolution: different cognitive systems converge on the same abstraction, when facing similar selection pressures and learning in similar environments. It is a less ambitious alternative to 'natural abstractions hypotheses' and, in my view, more likely to be true. Convergence may be real, useful, and empirically robust, while still being contingent and fragile under changes in architecture,
AI SafetyAlignment TheoryMechanistic InterpretabilityNatural Abstractions
Research LessWrong May 15

Mechanistic estimation for expectations of random products

By Jacob_Hilton

55 score
AI Analysis

Presents methods for mechanistic estimation of expectations of random products, applicable to problems like random halfspace intersections, random #3-SAT, and random permanents. The methods are competitive with sampling-based approaches and build on the 'matching sampling principle.'

We have developed some relatively general methods for mechanistic estimation competitive with sampling by studying problems that are expressible as expectations of random products. This includes several different estimation problems, such as random halfspace intersections, random #3-SAT and random permanents. In this post, we will give a high-level introduction to these methods before sharing some more detailed notes. This is intended as an interim technical update and will be relatively light o
Computational TheoryEstimation MethodsMathematics
Research LessWrong May 15

Announcing the Center for Shared AI Prosperity

By Dylan Matthews

52 score
AI Analysis

Announces a new DC-based policy organization focused on AI's economic impacts, led by David Shor, Stef Feldman, and others. The center targets taxation, income support, workforce policy, and industrial policy to address wealth concentration and labor displacement from advanced AI.

I wanted to share the launch of a project I've been working on with pollster David Shor, Obama/Biden veteran Stef Feldman, political strategist Morris Katz, Harvard historian Marc Aidinoff, and a few other folks*.The Center for Shared AI Prosperity is an attempt to force DC policy elites, particularly (given our team's backgrounds) liberals/progressives, to take the impending economic impacts of advanced AI more seriously. We do not think this is a normal economic shock. We are deeply uncertain
AI PolicyEconomic ImpactAI GovernanceLabor Displacement
Research LessWrong May 14

The hard core of alignment (is robustifying RL)

By Cole Wyeth

48 score
AI Analysis

Argues that the fundamental difficulty of alignment is making reinforcement learning robust - that across diverse alignment approaches, there is a common 'hard core' problem analogous to a complete problem in complexity theory. Claims most safety work fails to engage with this core challenge.

Most technical AI safety work that I read seems to miss the mark, failing to make any progress on the hard part of the problem. I think this is a common sentiment, but there's less agreement about what exactly the hard part is? Characterizing this more clearly might save a lot of time and better target the search for solutions. In this post I explain my model of why alignment is technically hard to achieve, setting aside the regulatory, competitive, and geopolitical challenges, the sheer incompe
AI SafetyAlignmentReinforcement LearningAI Theory
Research LessWrong May 15

Clarifying the Darwinian Honeymoon

By Elias Schmied

30 score
AI Analysis

Clarifies a previous post arguing that historical dominance in evolutionary competition doesn't evidence future dominance - using the analogy of chickens thriving under human domestication before being outcompeted. The core claim is that the 'outside view' of human progress provides less evidence against displacement than commonly assumed.

A few days ago, I published The Darwinian Honeymoon - why I am not as impressed by human progress as I used to be. To my gratification, it was quite well-received on Twitter, Substack and LessWrong.[1]However, in subsequent conversations I realized that I did not communicate my core point well enough, given how abstract it is. So I wanted to write a short (and somewhat sloppy and galaxy-brained) post explaining what I mean a bit more.Here’s what I am NOT saying:“Humans will face the same fate as
AI Existential RiskEvolutionary TheoryAI Forecasting
25 score
AI Analysis

Argues against AI existential risk using two principles: Copernicanism (if superintelligence were common, we'd observe cosmic evidence like Dyson spheres) and the straight-line extrapolation principle. Essentially a Fermi paradox-based argument against AI takeoff scenarios.

I have no good gears-level model of AI, and the expert views are all over the place (see AI Doc), so the only remaining argument is my physical intuition and a black-box view. Which, in this case, is based on two principles (best called rules of thumb): Copernicanism, where one assumes that humanity is not unique or special among the stars The "Law" of Straight Lines: slatestarcodex.com/2019/03/13/does-reality-drive...
AI Existential RiskFermi ParadoxAI Forecasting
20 score
AI Analysis

Argues that data quality in Africa is critically poor and that development organizations should prioritize funding technical assistance programs to improve statistical capacity. Claims that many data points from African countries are essentially fabricated, undermining evidence-based interventions.

The title for this post is inspired by: Forecasting is Way Overrated, and We Should Stop Funding It — LessWrongSummaryData quality in Africa is near-universally poor, especially at a sub-national level. Organisations and individuals who care about development, poverty alleviation and social welfare should fund measures and programmes that improve data quality via technical assistance. With good data, ‘mysteries’ of African (under)development can be better addressed, and more people can be lifted
Data QualityDevelopment EconomicsEffective Altruism
Research LessWrong May 15

Monthly Roundup #42: May 2026

By Zvi

18 score
AI Analysis

Zvi Moshowitz's monthly roundup covering a hantavirus outbreak on a cruise ship, predictions, technology advances, the Jones Act, and various other topics. A broad news digest with commentary rather than focused research.

At least we probably won’t have another pandemic. And we still have a partial Jones Act waiver. For now. Small victories. Table of Contents Hanta Hanta I Don’t Wanta. Bad News. Predictions Can Be Easy Even About The Future. Good Advice. The Efficient Market Hypothesis Is False. There Are Four Skills. While I Cannot Condone This. Good News, Everyone. For Your Entertainment. Gamers Gonna Game Game Game Game Game. The Spire Sleeps And So Shall I. I Was Promised Flying Self-Driving Cars. Government
News CommentaryForecastingPolicy
Research LessWrong May 15

MATS 9 Retrospective & Advice

By beyarkay

15 score
AI Analysis

A retrospective from a MATS (ML Alignment Theory Scholars) Season 9 alum describing the work ethic, research environment, and practical advice for future applicants. The author worked on Team Shard with Alex Turner and Alex Cloud from January to March 2026.

I couldn’t find a recent write-up from a MATS alum about what attending MATS was like, so this is the thing that I wish I had. I attended MATS from January to March 2026, on Team Shard with Alex Turner and Alex Cloud. It was a great time! Applications for MATS are basically on a rolling basis nowadays, and I can strongly recommend applying (to multiple streams) even if you think you’re not a great match.With that being said, there’s a lot I wish I knew going into MATS, so here’s a brain-dump of
AI Safety CommunityResearch ProgramsCareer Advice