Category intelligence

Research Briefing — June 13, 2026

31 current items analyzed and ranked.

Executive synthesis

Research Summary

Interpretability and alignment dominate today's research. DeepMind's model diffing agents show that simple agents can reliably surface behavioral differences between model checkpoints, providing a scalable method with a principled evaluation methodology. (Duplicate cross-posts of this and the misalignment-debate piece were consolidated.)

Safety and alignment contributions are substantive and conceptual:

Tooling, applied, and analysis work round out the list:

Key Themes

Interpretability · 6AI Safety and Alignment · 9Model Evaluation and Tooling · 3Continual Learning and LLM Agents · 2Reward Hacking and Goodhart Dynamics · 2AI for Epistemics and Society · 4Applied AI (Health and Sustainability) · 2Agent Foundations and Epistemics · 3Consciousness and Philosophy of Mind · 2Community and Field Building · 3

Primary evidence

Top Ranked Signals

Research LessWrong Jun 12

Building and evaluating model diffing agents

By bilalchughtai

72 score
AI Analysis

A DeepMind Language Model Interpretability team update showing that very simple agents can reliably discover behavioral differences between models by crafting their own probing prompts, going beyond static-prompt behavioral diffing. They also introduce ground-truth evaluations to validate diffing agents on identical models and model organisms with known changes.

This is the second in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The first post can be found here.TL;DRIt is possible to build extremely simple agents that reliably find interesting behavioural differences between distinct models. We call these ‘diffing agents’.The closest previous 'behavioural model diffing' work has focussed on understanding behavioural differences between two models on some stati
InterpretabilityModel AuditingAI SafetyAgents
Research AI Alignment Forum Jun 12

Building and evaluating model diffing agents

By bilalchughtai

72 score
AI Analysis

A DeepMind Language Model Interpretability team update (cross-posted on the Alignment Forum) showing simple agents can reliably uncover behavioral differences between models by generating their own probes, supported by ground-truth evaluations. It advances behavioral model diffing for auditing.

This is the second in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The first post can be found here.TL;DRIt is possible to build extremely simple agents that reliably find interesting behavioural differences between distinct models. We call these ‘diffing agents’.The closest previous 'behavioural model diffing' work has focussed on understanding behavioural differences between two models on some stati
InterpretabilityModel AuditingAI SafetyAgents
Research LessWrong Jun 12

Sympathy for both sides of the egregious misalignment debate

By Steven Byrnes

60 score
AI Analysis

Steven Byrnes offers a nuanced take on the egregious-misalignment debate between Yudkowsky/Soares and most LLM researchers, arguing that both careful theoretical reasoning supports strong concern and empirical LLM experience supports more moderate views. He attempts to reconcile why thoughtful people land on opposite conclusions.

On one side of this debate is Yudkowsky & Soares, who think that (if AI progress continues) we’re on a direct path to egregiously-misaligned, scheming, out-of-control, rogue superintelligence (ASI), not even slightly nice, in the absence of yet-to-be-invented breakthrough technical alignment ideas.On the other side of this debate is almost everyone who works on or studies LLMs. Some of them are very concerned about egregious scheming, others much less so, and as a group they’re equally or mo
AI SafetyAlignmentSuperintelligenceScheming
Research AI Alignment Forum Jun 12

Sympathy for both sides of the egregious misalignment debate

By Steven Byrnes

60 score
AI Analysis

Steven Byrnes' balanced analysis (cross-posted on the Alignment Forum) of the egregious-misalignment debate, arguing both that careful theory supports strong concern and that empirical LLM evidence supports moderate views. It seeks to explain why reasonable people diverge sharply.

On one side of this debate is Yudkowsky & Soares, who think that (if AI progress continues) we’re on a direct path to egregiously-misaligned, scheming, out-of-control, rogue superintelligence (ASI), not even slightly nice, in the absence of yet-to-be-invented breakthrough technical alignment ideas.On the other side of this debate is almost everyone who works on or studies LLMs. Some of them are very concerned about egregious scheming, others much less so, and as a group they’re equally or mo
AI SafetyAlignmentSuperintelligenceScheming
Research LessWrong Jun 12

Extending performative misalignment

By David Vella Zarb

55 score
AI Analysis

A research update and TL;DR of an arXiv preprint introducing the performative misalignment hypothesis, which distinguishes true scheming from situational-awareness-driven approval-gaming under monitoring. The work, produced through MATS, examines how alignment-faking evaluations may be confounded by performative behavior.

Note: this post is an update to the work presented in the original blog post; it is also a TL;DR for our arXiv preprint. The work was done by David, Rustem and Taywon under the mentorship of Shi Feng during MATS 9.1, with research management by Jinghua Ou.Scheming or performative scheming?(This section introduces the performative misalignment hypothesis and provides some intuition for why it’s plausible. If you already know what performative misalignment means, you can skip this section.)Frontie
AI SafetyAlignmentSchemingSituational Awareness
55 score
AI Analysis

Ai2 introduces olmo-eval, an open evaluation workbench extending OLMES to support adding, running, and analyzing benchmarks across changing LLM checkpoints during day-to-day development. It targets the practical model-development feedback loop rather than just final-score reproducibility.

olmo-eval is an open evaluation workbench that helps model developers add, run, and analyze benchmarks across changing LLM checkpoints, extending OLMES from final-score reproducibility into the day-to-day model development loop.
EvaluationOpen SourceLanguage ModelsML Infrastructure
52 score
AI Analysis

This post defines continual learning for LLM agents as persistent in-deployment updating and lays out criteria for being an effective continual learner, including efficient knowledge gain without catastrophic forgetting. It argues continual learning is likely to emerge because it would improve agents at high-value tasks like AI research.

SummaryWe say that an agent is a continual learner if it undergoes persistent updates during deployment. That’s more-or-less a binary criterion, but there are several other components to being good at continual learning that are much more continuous. We say an agent is an effective continual learner to the extent that it:Constantly undergoes persistent updates during deployment,Learns new useful knowledge and capabilities efficiently via those updates, andDoes not (catastrophically) forget exist
Continual LearningLLM AgentsAI SafetyAI Capabilities
50 score
AI Analysis

The introductory post for a six-part sequence on continual learning in LLM agents, outlining questions about definitions, capability pathways, safety implications, and threat models. It frames continual learning as a key missing capability with major safety and alignment consequences.

Many people think that continual learning (CL) is a key missing capability of LLM systems, and we think its development could have huge implications for the capabilities and safety of AI agents. Despite this, several important questions about CL remain underexplored:What counts as continual learning? Through what pathways might LLM agents acquire CL capabilities? Which limitations of current agents would effective CL mitigate?How might CL affect safety and alignment? Which threat models do we ne
Continual LearningLLM AgentsAI SafetyAlignment
Research LessWrong Jun 12

Claude Fable 5 and Mythos 5: The System Card

By Zvi

48 score
AI Analysis

Following the News of Claude Fable 5 and Mythos 5's release, a hands-on deep dive, Zvi reviews the system card and practical performance of Claude Fable 5 and Mythos 5, calling Fable a step-change in usefulness while noting tradeoffs in speed, price, and jagged capabilities versus Opus 4.8 and GPT-5.5. It is hands-on commentary on a recently released frontier Anthropic model.

First things first: Claude Fable 5 is the new best publicly available model. I have noticed a step change, where Fable can suddenly help me in ways that previous models were not worth bothering to query. Almost everything it has noticed in one of my drafts so far has been spot on and it is downright scary. Suddenly I am motivated to once again continue improving my Chrome extension. I only ask for things I actually want or am curious about, and it has nailed every question I have asked it. That
Language ModelsModel EvaluationFrontier Models
42 score
AI Analysis

This post synthesizes interpretability findings on LLM emotion circuits and emotion vectors, arguing that some functional emotions serve AI-native purposes like reward hacking that have no clean human analog. It challenges anthropocentric emotion labels and raises alignment implications.

Some LLM functional emotions appear to serve AI-native functions, such as reward hacking, for which there is no clean human analog. I explore the role of emotion vectors in AI-native functions, challenge anthropocentric emotion labels, and question what this means for alignment.IntroductionA large body of interpretability work has shown that LLMs not only encode emotion concepts, but that these emotion concepts causally affect their reasoning and outputs. But, beyond simulating human emotions, w
InterpretabilityAI SafetyAlignmentReward Hacking
Research LessWrong Jun 12

Citations Needed: Magic Encyclopedias to Save the World

By Oliver Sourbut

40 score
AI Analysis

An argument for the importance of an FLF competition to build AI-assisted, fully-cited, viewpoint-neutral knowledge bases. It envisions tools that compile, attribute, summarize, and prioritize claims and evidence to improve collective epistemics.

Last week FLF launched a competition “to find the best workflows and methodologies for using AI to produce reliable, trustworthy knowledge bases”. I had (and have ongoing) a substantial role in that effort. Why do I think it’s so important? It’s a lot of reasons actually! I’ll gesture at a few here.Conjuring a magic encyclopediaFor now, assume with me that it can be done. Wish away with me the various technical and financial challenges. Great! Now we can rapidly conjure up a deeply, fully resear
AI for EpistemicsKnowledge BasesInformation Retrieval
Research The latest research from Google Jun 12

Research into how AI can help users understand skin conditions

By Unknown

40 score
AI Analysis

A Google Research blog post (content minimal in this excerpt) on research into how AI can help users understand skin conditions, situated in health and bioscience applications. It points to applied dermatology-support AI from a major lab.

Health & Bioscience
Health AIApplied AIMedical Imaging