Category intelligence

Research Briefing — May 24, 2026

17 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's posts emphasize AI safety research, governance analysis, and foundational philosophy, with limited frontier model research.

Safety & Alignment Research:

Governance & Strategic Analysis:

Philosophy & Foundations: Posts sketch unified frameworks for Bayesian priors and anthropics, argue Boltzmann brain and Doomsday puzzles need no explanation, and critique 'single genius' framings of AGI cooperation dynamics.

Key Themes

AI Safety & Alignment · 6AI Governance & Policy · 4Philosophy & Decision Theory · 4Rationalist Lifestyle/Commentary · 4

Primary evidence

Top Ranked Signals

70 score
AI Analysis

Owain Evans provides a primer and reading list on out-of-context reasoning (OOCR) in LLMs - cases where models combine facts in the forward pass without verbalized chain-of-thought. Highly relevant to alignment given implications for hidden reasoning and deceptive capabilities.

Out-of-context reasoning (OOCR) is a concept relevant to LLM generalization and AI alignment. Also available as a PDF. Contents What is OOCR? Examples Papers Videos What is out-of-context reasoning for LLMs? It's when an LLM reaches a conclusion that requires non-trivial reasoning but the reasoning is not present in the context window. The reasoning could instead take place in the forward pass or during the training process. The name ("out-of-context reasoning") is chosen to contrast with in-con
AI AlignmentLLM GeneralizationInterpretability
55 score
AI Analysis

Translation and analysis of a January 2025 PLA Daily article by Chinese military authors on AGI's implications for warfare, with framing context arguing China is not actually racing for frontier AGI. Useful primary source for AI race discourse.

Source“Reflections on Warfare Brought by AGI” (AGI带来的战争思考)Source: PLA Daily (解放军报)Date: January 21, 2025Authors: Rong Ming (荣明), Hu Xiaofeng (胡晓峰)IntroductionPlease feel free to skip to the translation, about halfway down, though I would recommend reading the sections “On the source” and "On the Authors" just above it too.In November 2024, the U.S.-China Economic and Security Review Commission recommended that “Congress establish and fund a Manhattan Project-like program dedicated to racing to a
AI GeopoliticsAI PolicyMilitary AI
55 score
AI Analysis

Tests whether LLMs refuse uplift requests on mirror life - a real emerging biothreat not yet officially classified as WMD/CBRN. Examines the gap between safety training and unclassified novel threats.

[Cross-posted from On Failure States. This is Part 1 of an independent AI safety research series examining LLM safety behavior on unclassified emerging threats.]Can an LLM refuse a harmful uplift request when the topic in question hasn’t been identified as dangerous yet? In 2022, mirror RNA polymerase was actually created, a key step towards the creation of mirror life, and in 2024 the scientific community warned against any further research on it.[1][2] Having said that, mirror life is not curr
AI SafetyBiosecurityLLM Evaluation
Research LessWrong May 22

How should we update on AI-enabled coups post-Mythos?

By boygirlseating

50 score
AI Analysis

Analyzes how Anthropic's Claude-Mythos-Preview (a model deemed too dangerous, with major cyber capabilities) should update beliefs about AI-enabled coup risks. Argues Mythos lowers minimum coalition size for targeted disruption while concentrating decision-power in private actors.

Last month, Anthropic developed Claude Mythos, a model they considered too dangerous for public release.As per Anthropic (and via testing from AISI), we know that Mythos:Found thousands of previously unknown vulnerabilities in every major operating system and browser.Surpasses the coding capabilities of all but the most skilled humans.Exposed a 25+-year-old flaw in the world’s most secure operating system that would let it crash essential infrastructure.There’s a great write up from 80,000 Hours
AI SafetyAI GovernanceCyber Risk
Research LessWrong May 22

Looking for backdoors in Jane Street LLMs

By Cipolla

50 score
AI Analysis

Hands-on writeup of attempting Jane Street's LLM backdoor detection challenge using white-box methods on fine-tuned Qwen2.5-7B and DeepSeek-V3 models. Reports partial success after activation/prompting approaches failed.

I am going to talk about my experience in the Jane Street LLM backdoor challenge. I am sharing partial results. I managed to crack some of the models using white-box methods, after the activation/prompting approach didn't pan out. Happy to discuss better or more promising approaches.IntroductionA few months ago a Dwarkesh Patel podcast episode advertised a Jane Street backdoor challenge:We've trained backdoors into three language models.On the surface, they behave like ordinary conversational mo
AI SafetyBackdoorsInterpretabilityRed-teaming
Research LessWrong May 23

Probabilities are not the right concept

By David Matolcsi

35 score
AI Analysis

First post in a planned sequence sketching a unified framework for Bayesian priors, probability foundations, anthropics, and infinite ethics. Synthesizes existing work from Christiano, Garrabrant, Carlsmith and others.

IntroductionThis sequence is an attempt to sketch a unified framework for several interconnected questions: Where do Bayesian priors come from? What even are probabilities? How should we deal with infinite ethics? What's going on with anthropics? I hope to lay out both some of the existing answers and my own preferred synthesis.[1]I understand that many people have already thought about these questions, and I have only read portions of the existing literature. I think most of what I will write h
Decision TheoryProbabilityAnthropics
Research LessWrong May 23

Your Left Brain Doesn't Trade With Your Right

By Alexander Gietelink Oldenziel

30 score
AI Analysis

Critique of 'single genius model' framing of AGI, arguing economists miss that AGI may transcend the cooperation/specialization paradigm by integrating multiple perspectives in one mind. Conceptual essay on AGI nature.

[see also Four Ways Learning Economics makes you people dumber future AI]This is a tweet by Seb Krier that caught my eye. The exact person and exact points are incidental. It illustrates what to is a flaw in many 'economics' frames on AI. Expecting a model to do all the work, solve everything, come up with new innovations etc is probably not right. This was kinda the implicit assumption behind *some* interpretations of capabilities progress. The ‘single genius model’ overlooks the fact that
AGIEconomics of AI
25 score
AI Analysis

An informal LessWrong post describing a 'vibe coded' experiment to use LLM agents to iterate on minimal agent-based models (Sugarscape), with a Karpathy-style ablation harness. Explores LLM-driven autoresearch but is preliminary and not rigorous.

IntroAgent based models (ABMs) are notorious for being over-parameterized. Researchers can get caught in loops effectively doing manual p-hacking adding more and more rules and variables to their funny world of little guys in a grid. My problem is that ABMs can be far too complex, to truly understand the emergent phenomena (in my opinion) we must find the simplest structure that results in their emergence. With this in mind, I vibe coded and iterated on a small experiment: a harness to allow for
LLM AgentsAutoresearchAgent-Based Models
Research LessWrong May 23

Boltzmann brains, like Doomsday, require no explaining

By Steffee

25 score
AI Analysis

Philosophical post arguing that the Doomsday Argument and Boltzmann brain anthropic puzzles don't actually require explanation, challenging Yudkowsky's recent post. Engages with anthropic reasoning.

Brothers and sisters I have none, but that man's father is my father's son. Who am I?— ancient riddleIn Eliezer Yudkowsky’s post this week, he writes: “Our current experience -- your own experience, at this very moment, of seeing ordered letters on a screen -- therefore seems to provide overwhelming anthropic evidence against any model of reality or physics which would imply that most brains are Boltzmann brains.”I hope it’s fair for me to roughly present this line of reasoning like so:If Boltzm
AnthropicsPhilosophy
Research LessWrong May 22

A political movement will save us from extinction

By rohantohab

25 score
AI Analysis

Founder of 'Sapiens First' argues AI safety needs a populist political movement to build societal capacity for governance. Activism-oriented post citing the (fictional in reality but in-world) Mythos model.

Epistemic status: I have spent the better part of 3.5 months full-time thinking about AI safety politics. I used to work at Constellation, and have followed rationalism for about a decade. Before working in AI safety, I was deeply involved in the animal rights movement, which may inform my views.This post includes claims about AI politics. I'm very interested in folks agreement/disagreement. For example, if you disagree, which statement do you disagree with, and why? I invite you to steel-man, i
AI GovernanceAI PolicyAI Safety
Research LessWrong May 22

The Leaky AI Safety Pipeline

By Nikhil Kalidasu

25 score
AI Analysis

Reflection on becoming a productive AI safety researcher, citing acceptance of an independent paper at the SiMLA workshop at ACNS 2026. Argues field should optimize the conversion pipeline from program graduates to peer-reviewed researchers.

My first AI security paper as an independent researcher (with one other independent collaborator) was just accepted to the Security in Machine Learning Applications workshop at ACNS 2026. This was an 8-month process: I spent 2 weeks convincing myself that what I was trying to do was possible, 4 months building, evaluating, and collecting results, and another 4 months writing the paper (and making strategic mistakes) before submitting.My paper was not radical or paradigm-shifting by any means, bu
AI SafetyResearch Community
Research LessWrong May 22

Capitalism is only the first of our problems

By Ian Matson

15 score
AI Analysis

Argues AI alignment must precede economic restructuring discussions about AI's threat to capitalism. Commentary essay.

I've been thinking a lot about the impact of AI on the job market recently, and as a result, I ended up reading quite a few papers/articles on the threat that AI poses to capitalism. However, the more I read, the more it seemed to me that the authors of these papers were missing a key assumption in their proposed solutions to this issue - that every positive outcome, for capitalism for for humanity, is predicated first on the alignment of AI to the best interests of humankind.The Threat to Capit
AI PolicyAI Alignment