Category intelligence

Research Briefing — January 20, 2026

16 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research concentrates heavily on alignment techniques and safety evaluation. A survey on alignment pretraining synthesizes evidence that training LLMs on data depicting well-behaved AI during pretraining substantially reduces misalignment—potentially offering a scalable, proactive safety approach.

  • Coup probes testing demonstrates few-shot linear classifiers can detect scheming behavior from model activations, with empirical results on off-policy training data
  • Silent Agreement Evaluation provides first empirical measurement of Schelling coordination in LLMs—whether isolated instances converge on shared choices without communication
  • Framework for AI-delegated safety research identifies key dimensions: epistemic cursedness, parallelizability, and short-horizon suitability
  • Strategic analysis examines whether LLM alignment work transfers to non-LLM takeover-capable systems

Methodological contributions include a critique of METR-HRS timelines forecasting, arguing the 'd' parameter conflates task difficulty with sequence length. Governance-oriented work sketches positive AI transition scenarios co-authored with Claude Opus 4.5.

Key Themes

AI Safety and Alignment · 6AI Capabilities and Evaluation · 3AI Governance and Futures · 2Decision Theory and Philosophy · 4

Primary evidence

Top Ranked Signals

75 score
AI Analysis

Survey of 'alignment pretraining' research showing that training LLMs on data depicting AI behaving well during pretraining dramatically reduces misalignment, and this persists through post-training. Claims major labs are now interested in this approach.

Alignment Pretraining Shows PromiseTL;DR: A new paper shows that pretraining language models on data about AI behaving well dramatically reduces misaligned behavior, and this effect persists through post-training. The major labs appear to be taking notice. It’s now the third paper on this idea, and excitement seems to be building.How We Got Here(This is a survey/reading list, and doubtless omits some due credit and useful material — please suggest additions in the comments, so I can update it. O
AI SafetyAlignmentLanguage ModelsPretraining
Research LessWrong Jan 19

Testing few-shot coup probes

By Joey Marcellino

70 score
AI Analysis

Implements and tests linear classifiers (coup probes) trained on AI activations to detect scheming behavior. Tests whether off-policy training data can bootstrap detection that improves with real examples. First empirical test of this proposed technique.

I implemented (what I think is) a simple version of the experiment proposed in [1]. This is a quick writeup of the results, plus a rehash of the general idea to make sure I’ve actually understood it.ConceptWe’d like to be able to monitor our AIs to make sure they’re not thinking bad thoughts (scheming, plotting to escape/take over, etc). One cheap way to to do this is with linear classifiers trained on the AI’s activations, but a good training dataset is likely going to be hard to come by, since
AI SafetyInterpretabilityAlignmentAI Monitoring
Research LessWrong Jan 19

Silent Agreement Evaluation

By Graeme Ford

68 score
AI Analysis

First empirical study measuring Schelling coordination in LLMs - whether two model instances independently choose the same option without communication. Frontier models failed at chance without reasoning; thinking models succeeded on word comparisons.

Measuring out-of-context Schelling coordination capabilities in large language models.OverviewThis is my first foray into AI safety research, and is primarily exploratory. I present these findings with all humility, make no strong claims, and hope there is some benefit to others. I certainly learned a great deal in the process, and hope to learn more from any comments or criticism—all very welcome. A version of this article with some simple explanatory animations, less grainy graphs, and updates
AI CapabilitiesMulti-Agent SystemsEvaluationAI Safety
Research LessWrong Jan 19

Desiderata of good problems to hand off to AIs

By Jozdien

65 score
AI Analysis

Framework identifying key dimensions for which AI safety problems to delegate to AI systems: epistemic cursedness, parallelizability, short-horizon sub-problems, speed of ASI alignment progress, and legibility to labs.

Many technical AI safety plans involve building automated alignment researchers to improve our ability to solve the alignment problem. Safety plans from AI labs revolve around this as a first line of defence (e.g. OpenAI, DeepMind, Anthropic); research directions outside labs also often hope for greatly increased acceleration from AI labor (e.g. UK AISI, Paul Christiano).I think it’s plausible that a meaningful chunk of the variance in how well the future goes lies in how we handle this handoff,
AI SafetyAlignmentResearch StrategyAI Automation
62 score
AI Analysis

Analyzes whether LLM alignment research transfers to non-LLM AIs via two mechanisms: direct transfer (reusing evaluations, model organisms) and indirect transfer (using aligned LLMs to oversee non-LLMs). Argues surprisingly much research may transfer directly.

Many people believe that the first AI capable of taking over would be quite different from the LLMs of today. Suppose this is true—does prosaic alignment research on LLMs still reduce x-risk? I believe advances in LLM alignment research reduce x-risk even if future AIs are different. I’ll call these “non-LLM AIs.” In this post, I explore two mechanisms for LLM alignment research to reduce x-risk:Direct transfer: We can directly apply the research to non-LLM AIs—for example, reusing behavioral ev
AI SafetyAlignmentResearch StrategyX-Risk
Research LessWrong Jan 19

AGI both does and doesn't have an infinite time horizon

By Sean Herrington

55 score
AI Analysis

Critiques AI timelines forecasting by arguing that METR-HRS tasks conflate difficulty with sequence length. The key parameter 'd' in forecasting models depends on whether intelligence or consistency is the bottleneck, dramatically changing extrapolations.

TLDR Long time horizon METR-HRS tasks are both more difficult and sequentially longer than short tasksThe resulting benchmark is therefore measuring both the ability to complete difficult tasks and consistency in its abilities over long time frames.Depending on whether you think intelligence or consistency is the bottleneck, your extrapolated time horizons change dramaticallyIn particular, I expect people who see intelligence as the bottleneck to extrapolate an infinite (or extremely large)
AI CapabilitiesAGI TimelinesBenchmarksForecasting
Research LessWrong Jan 19

Gradual Paths to Collective Flourishing

By Nora_Ammann

52 score
AI Analysis

Positive scenario for AI transition co-written with Claude, identifying failure modes (unilateral domination, competitive erosion) and sketching a path through gradual takeoff and multipolarity toward collective flourishing.

by Nora Ammann & Claude Opus 4.5Setting the stageThere aren't many detailed stories about how things could go well with AI.[1] So I'm about to tell you one. This is an attempt to articulate a path, through the AI transition, to collective flourishing. What makes endgame sketches like this useful is that they need to be constrained. They need to be coherent with your best guess of how the world works, and earnestly engaged with good-faith articulations of risks and failure mode
AI GovernanceAI SafetyAI FuturesExistential Risk
32 score
AI Analysis

A philosophical critique arguing that even if we solved computational intractability in longtermist expected value calculations, we still couldn't make definitive decisions. Extends Kinney's 2022 work on why longtermist effective altruism may not be action-guiding.

Dnnn Uunnn, nnn nnn nnn nuh nuh nuh nuh, dnnn unnn nnn nnn nnn nuh nuh nuh NAH (Tears for Fears) I was reading David Kinney’s interesting work from 2022 “Longtermism and Computational Complexity” in which he argues that longtermist effective altruism is not action-guiding because calculating the expected utility of events in the far future is computationally intractable. The crux of his argument is that longtermist reasoning requires probabilistic inference in causal models (Bayesian n
Decision TheoryEffective AltruismPhilosophy
Research LessWrong Jan 18

Five Theses on AI Art

By jenn

25 score
AI Analysis

Essay comparing AI art skepticism to Virginia Woolf's 1926 skepticism about cinema, arguing new mediums initially appear parasitic but develop unique capabilities. Defends AI art's potential through historical analogy.

1. We've Been On This Ride BeforeVirginia Woolf, writing at the dawn of cinema (1926), expresses doubt about whether or not this new medium has any legs:"Anna [Karenina] falls in love with Vronsky” – that is to say, the lady in black velvet falls into the arms of a gentleman in uniform and they kiss with enormous succulence, great deliberation, and infinite gesticulation, on a sofa in an extremely well-appointed library, while a gardener incidentally mows the lawn. So we lurch and lumber through
AI ArtCultural CommentaryMedia Theory
Research LessWrong Jan 19

All (Non-Trivial) Decisions Are Undecidable

By (M)ason

22 score
AI Analysis

Argues that finding optimal decision-making algorithms is undecidable by invoking the halting problem, claiming any such algorithm must encode an ideal halting point which cannot be determined.

A core tenet of rationalism is that for any given set of known information, a span of possible decisions, and some utility function, there exists a decision-making algorithm that will return the optimal outcome. A short argument, given below, will demonstrate (given relatively minimal assumptions) that finding any such algorithm is an undecidable problem.We will take as given three axioms:Running any decision-making algorithm requires a non-negligible amount of time, and incurs a non-negligible
Decision TheoryComputability TheoryPhilosophy
Research LessWrong Jan 19

"Lemurian Time War" by Ccru

By Nathan Delisle

18 score
AI Analysis

Discussion of Nick Land and the Cybernetic Cultural Research Unit, drawing parallels between their 'hyperstition' concept and LessWrong/Rationalist thinking as responses to similar cultural conditions.

"In the hyperstitional model Kaye outlined, fiction is not opposed to the real. Rather, reality is understood to be composed of fictions—consistent semiotic terrains that condition perceptual, affective and behavioral responses." - PdfThe Cybernetic Cultural Research Unit, founded by Nick Land, was a group of social theorists and philosophers at the University of Warwick who wrote about the internet, society, and culture. I understand that LessWrongers are generally ontologically skeptical of th
PhilosophyIntellectual HistoryCulture
Research LessWrong Jan 19

What can Kickstarter teach us about goal completion?

By Elijah

15 score
AI Analysis

Analysis of Kickstarter video game project completion rates, finding previous estimates of ~9% failure rate were methodologically flawed. Provides new data from 2014 and 2022 projects.

I, like many others, struggle with sticking to my goals. I was interested in analyzing data relevant to the topic and thought the crowdfunding platform Kickstarter might be an interesting place to look, as I was aware that not every funded Kickstarter delivered a product.I focused on video games that were successfully funded. I used a large dataset containing information about Kickstarter projects,[1] from which I randomly selected[2] fully-funded video game Kickstarters from 2014 and
Data AnalysisGoal Completion