Category intelligence

Research Briefing — January 11, 2026

13 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research centers on AI safety fundamentals and alignment tractability debates. A substantive technical argument against continuous chain-of-thought (neuralese) challenges OpenAI's research direction, claiming discrete tokens are architecturally necessary rather than bandwidth limitations.

Supporting work includes the False Confidence Theorem applied to Bayesian reasoning, a conceptual framework distinguishing superagency from superintelligence, and practical tooling applying PageRank to identify high-signal voices in AI discourse networks.

Key Themes

AI Safety & Alignment · 6Language Models & Architecture · 1Forecasting & Risk Assessment · 2Epistemics & Rationality · 4AI Community & Education · 2

Primary evidence

Top Ranked Signals

Research LessWrong Jan 10

The Case Against Continuous Chain-of-Thought (Neuralese)

By RobinHa

68 score
AI Analysis
Argues against continuous chain-of-thought ('neuralese') approaches, claiming that discrete tokens aren't just bandwidth limitations but actually necessary for error correction. Continuous latent representations would accumulate noise across reasoning steps, while discretization identifies and corrects errors.
Main thesis: Discrete token vocabularies don't lose information so much as they allow information to be retained in the first place. By removing minor noise and singling out major noise, errors become identifiable and therefore correctable, which continuous latent representations fundamentally cannot offer.The Bandwidth Intuition (And Why It's Incomplete)One of the most elementary ideas connected to neuralese is increasing bandwidth. After the tireless mountains of computation called a forward p
Language ModelsArchitectureChain-of-ThoughtNeural Network Design
62 score
AI Analysis
Provides a learning-theoretic analysis of how efficiently we can train AI policies or activation monitors to detect and remove bad behaviors like sandbagging during safety research. The post explores sample complexity bounds for both direct policy training and monitor-based approaches to catching deceptive AI actions.
I'm worried about AI models intentionally doing bad things, like sandbagging when doing safety research. In the regime where the AI has to do many of these bad actions in order to cause an unacceptable outcome, we have some hope of identifying examples of the AI doing the bad action (or at least having some signal at distinguishing bad actions from good ones). Given such a signal we could: Directly train the policy to not perform bad actions. Train activation monitors to detect bad actions. Thes
AI SafetyAlignmentMachine Learning TheoryDeceptive Alignment
Research LessWrong Jan 9

AI Incident Forecasting

By cluebbers

58 score
AI Analysis
Hackathon-winning project that trained statistical models on the AI Incidents Database, forecasting 6-11x increase in AI-related incidents over five years, particularly in misuse, misinformation, and system safety categories.
I'm excited to share that my team and I won 1st place out of 35+ project submissions in the AI Forecasting Hackathon hosted by Apart Research and BlueDot Impact!We trained statistical models on the AI Incidents Database and predicted that AI-related incidents could increase by 6-11x within the next five years, particularly in misuse, misinformation, and system safety issues. This post does not aim to prescribe specific policy interventions. Instead, it presents these forecasts as evidence to hel
AI SafetyForecastingAI RiskAI Incidents
55 score
AI Analysis
Argues against optimistic views (citing Evan Hubinger) that alignment might be 'steam engine difficulty' - pointing out that steam engines took 70+ years from patent to practical vehicles. Even 'easy' alignment could fail if we lack sufficient time or coordination before dangerous capabilities emerge.
Cross-posted from my website. You may have seen this graph from Chris Olah illustrating a range of views on the difficulty of aligning superintelligent AI: Evan Hubinger, an alignment team lead at Anthropic, says: If the only thing that we have to do to solve alignment is train away easily detectable behavioral issues...then we are very much in the trivial/steam engine world. We could still fail, even in that world—and it’d be particularly embarrassing to fail that way; we should definitely make
AI SafetyAI PolicyAlignmentExistential Risk
Research LessWrong Jan 10

The false confidence theorem and Bayesian reasoning

By viking_math

52 score
AI Analysis
Introduces the False Confidence Theorem to LessWrong, arguing it explains why strong Bayesian arguments can feel intuitively wrong. Uses satellite conjunction analysis as exposition and suggests this theorem underlies errors in debates like Rootclaim's lab-leak analysis.
A little backgroundI first heard about the False Confidence Theorem (FCT) a number of years ago, although at the time I did not understand why it was meaningful. I later returned to it, and the second time around, with a little more experience (and finding a more useful exposition), its importance was much easier to grasp. I now believe that this result is incredibly central to the use of Bayesian reasoning in a wide range of practical contexts, and yet seems to not be very well known (I was not
EpistemicsBayesian ReasoningRationality
48 score
AI Analysis
Applies PageRank algorithm to Twitter's AI discourse network to identify 'important' and 'underrated' people to follow. Uses the intuition that if someone important follows few accounts, those accounts are likely high-signal. Includes LLM-based quality assessment.
Cross post, adapted for LessWrongSeveral challenges add friction to finding high signal people and literature:High status may negatively impact signal.Exploration can only be done at the edges of my network, e.g. Twitter thread interactions or recommended people to follow, bottlenecked by I don’t know what I don’t know.Recommendations naturally bias toward popular people.Even recommended people from a curated following list may be important but low signal, e.g. Sam Altman’s priority is promoting
Social Network AnalysisInformation DiscoveryAI Community
Research LessWrong Jan 10

Possible Principles of Superagency

By Mariven

45 score
AI Analysis
Proposes a conceptual framework for 'superagents' - actors achieving goals with greater efficiency than humans, likely consisting of human-AI teams before pure AI systems. Outlines principles including directedness and other properties that enable superagency.
Prior to the era of superintelligent actors, we’re likely to see a brief era of superagentic actors—actors who are capable of setting and achieving goals in the pursuit of a given end with significantly greater efficiency and reliability than any single human. Superagents may in certain restricted senses act superintelligently—see principles 8, 9—but this isn’t strictly necessary. A superagent may be constructed from a well-scaffolded cluster of artificial intelligences, but in the near-term it’
AI CapabilitiesAgent FoundationsHuman-AI Collaboration
40 score
AI Analysis
Proposes restructuring ARENA (AI safety training program) from contained exercises to four one-week research sprints, arguing the real bottleneck is research engineering skills rather than conceptual knowledge. Claims current ARENA primarily functions as a signaling mechanism.
TLDRI propose restructuring the current ARENA program, which primarily focuses on contained exercises, into a more scalable and research-engineering-focused model consisting of four one-week research sprints preceded by a dedicated "Week Zero" of fundamental research engineering training. The primary reasons are:The bottleneck for creating good AI safety researchers isn't the kind of knowledge contained in the ARENA notebooks, but the hands-on research engineering and research skills involved in
AI SafetyEducationCommunity Building
Research LessWrong Jan 9

What do we mean by "impossible"?

By Sniffnoy

35 score
AI Analysis
Reposted taxonomy of different meanings of 'impossible' - from logical contradictions (Level 0) through physical impossibility to merely difficult. Aims to clarify confused discussions where people use 'impossible' to mean different things.
(I'm reposting this here from an old Dreamwidth post of mine, since I've seen people reference it occasionally and figure it would be easier to find here.) So people throw around the word "impossible" a lot, but oftentimes they actually mean different things by it. (I'm assuming here we're talking about real-world discussions rather than mathematical discussions, where things are clearer.) I thought I'd create a list of different things that people mean by "impossible", in the hopes that it migh
PhilosophyEpistemicsRationalitySemantics
30 score
AI Analysis
Personal reflection on 'moral-epistemic scrupulosity' - the compulsive pursuit of certainty and correct thinking across frameworks like rationality and Tibetan Buddhism. Describes how this pattern manifests as intolerance of uncertainty and moralizing 'correct' reasoning.
Crossposted from substack.com/home/post/p-183478095 Epistemic status: Personal experience with a particular failure mode of reasoning and introspection that seems to appear within different philosophical frameworks (discussed here are rationality and Tibetan Buddhism), involving intolerance of felt uncertainty, over-indexing on epistemic rigour, compulsive questioning of commitments, and moralisation of "correct" thinking itself. If you do this correctly, you’ll be safe from er
EpistemicsRationalityPsychologyMetacognition
25 score
AI Analysis
Asks for strong arguments against acausal extortion (Roko's basilisk variants) being effective, noting that standard responses about precommitment have counter-responses. Seeks a clear, irrefutable reason not to worry about this class of decision-theoretic problems.
The topic of acausal extortion (particularly variants of Roko's basilisk) is sometimes mentioned and often dismissed with reference to something like the fact that an agent could simply precommit not to give in to blackmail. These responses themselves have responses, and it is not completely clear that at the end of the chain of responses there is a well defined, irrefutable reason not to worry about acausal extortion, or at least not to continue to do so once you have contemplated it. My questi
Decision TheoryAI SafetyAcausal Reasoning
Research LessWrong Jan 10

How Humanity Wins

By Wes R

18 score
AI Analysis
An emotionally-charged manifesto claiming to present a solution for humanity's existential challenges, referencing UN expert validation but providing no technical details or concrete proposals in the visible excerpt.
This isn't the full idea, but it basically is There’s a real chance that in the next 10 years we’ll all be dead because this year we didn’t get our act together. For the past 2 years, I’ve been researching & working on risks to Humanity's future, & this paper is about How Humanity Wins.There's so many issues out there, so many that could kill millions,So many that DO kill millions,Did you know there are trains carrying more explosives than Hiroshima travelling around the United
Existential RiskAI Governance