Category intelligence

Research Briefing — January 24, 2026

19 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research spans AI governance, safety evaluation, and foundational alignment theory. Peer-reviewed policy work proposes emergency response measures for catastrophic AI risk, specifically targeting gaps in Chinese AI regulation and deployment safety.

Meta-science initiatives propose systematic replication teams. Interpretability research examines attention sinks and the dark subspace where transformers store non-interpretable signals. Steven Byrnes releases v3 of his 225-page brain-like AGI safety resource.

Key Themes

AI Safety & Alignment · 8AI Governance & Policy · 2Interpretability & Mechanistic Understanding · 3Meta-Science & Research Methods · 2AI Consciousness & Ethics · 3Non-AI Content · 5

Primary evidence

Top Ranked Signals

Research LessWrong Jan 23

Emergency Response Measures for Catastrophic AI Risk

By MKodama

72 score
AI Analysis

Presents a paper on Chinese AI regulation and safety measures, arguing Chinese AI companies (like DeepSeek) lack adequate safety testing before deployment. Proposes emergency response frameworks and was presented at NeurIPS 2025 Workshop on Regulatable ML.

I have written a paper on Chinese domestic AI regulation with coauthors James Zhang, Zongze Wu, Michael Chen, Yue Zhu, and Geng Hong. It was presented recently at NeurIPS 2025's Workshop on Regulatable ML, and it may be found on ArXiv and SSRN.Here I'll explain what I take to be the key ideas of the paper in a more casual style. I am speaking only for myself in this post, and not for any of my coauthors.Thanks to James for creating this poster.The top US AI companies have better capabilities tha
AI GovernanceAI SafetyPolicyInternational AI Regulation
Research LessWrong Jan 23

Eliciting base models with simple unsupervised techniques

By Callum Canavan

68 score
AI Analysis

Empirical research testing simple unsupervised elicitation methods against the Internal Coherence Maximization (ICM) algorithm. Finds that few-shot prompts with random labels recover 53-93% of supervised performance, and identifies bootstrapping as ICM's most valuable component.

Authors: Aditya Shrivastava*, Allison Qi*, Callum Canavan*, Tianyi Alex Qiu, Jonathan Michala, Fabien Roger(*Equal contributions, reverse alphabetical)Wen et al. introduced the internal coherence maximization (ICM) algorithm for unsupervised elicitation of base models. They showed that for several datasets, training a base model on labels generated by their algorithm gives similar test accuracy to training on golden labels. To understand which aspects of ICM are most useful, we ran a couple of s
Base ModelsUnsupervised LearningElicitationLanguage Models
Research LessWrong Jan 23

A Framework for Eval Awareness

By LAThomson

65 score
AI Analysis

Proposes a conceptual framework for 'evaluation awareness'—when LLMs infer they're being evaluated and potentially behave differently. Introduces concepts like leveraging model uncertainty about eval type and awareness-robust consistency.

In this post, we offer a conceptual framework for evaluation awareness. This is designed to clarify the different ways in which models can respond to evaluations. Some key ideas we introduce through the lens of our framework include leveraging model uncertainty about eval type and awareness-robust consistency. We hope this framework helps to delineate the existing research directions and inspire future work.This work was done in collaboration with Jasmine Li in the first two weeks of MATS 9.0 un
AI SafetyEvaluationsDeceptionModel Behavior
Research LessWrong Jan 23

Digital Consciousness Model Results and Key Takeaways

By arvomm

62 score
AI Analysis

Introduces the Digital Consciousness Model (DCM), a probabilistic framework for assessing AI consciousness that incorporates multiple theories rather than assuming one. Presents initial results comparing different AI systems and biological organisms.

Introduction to the Digital Consciousness Model (DCM)Artificially intelligent systems, especially large language models (LLMs) used by almost 50% of the adult US population, have become remarkably sophisticated. They hold conversations, write essays, and seem to understand context in ways that surprise even their creators. This raises a crucial question: Are we creating systems that are conscious?The Digital Consciousness Model (DCM) is a first attempt to assess the evidence for consciousness in
AI ConsciousnessAI EthicsPhilosophy of MindMeasurement
Research LessWrong Jan 23

Paying attention to Attention Sinks

By Mitali M

58 score
AI Analysis

Summarizes research on 'attention sinks' and the 'dark subspace' in transformers—regions where models dump excess attention and store signals not intended for output. Removing this dark tail causes performance degradation, suggesting it serves a functional mechanical role.

I recently read Spectral Filters, Dark Signals, and Attention Sinks, an interesting paper on discovering where excess attention in transformers is dumped. Researcher found that transformers contain a "Dark Subspace" to store information that isn't intended for the output layer. The attention sink concept is a specific manifestation of this, where the model learns to dump the remaining attention from the Softmax into the first ([BOS]) token.The authors used spectral filters to decompose the resid
InterpretabilityTransformer ArchitectureMechanistic Understanding
Research LessWrong Jan 23

Principles for Meta-Science and AI Safety Replications

By zroe1

58 score
AI Analysis

Proposes founding a team dedicated to replicating AI safety research, outlining principles: meta-science shouldn't vindicate preexisting beliefs, selection of papers should be principled, and the field needs systematic verification given high stakes.

If we get AI safety research wrong, we may not get a second chance. But despite the stakes being so high, there has been no effort to systematically review and verify empirical AI safety papers. I would like to change that.Today I sent in funding applications to found a team of researchers dedicated to replicating AI safety work. But what exactly should we aim to accomplish? What should AI safety replications even look like? After 1-2 months of consideration and 50+ hours of conversation, t
Meta-ScienceAI SafetyReplicationResearch Methods
Research LessWrong Jan 23

New version of “Intro to Brain-Like-AGI Safety”

By Steven Byrnes

55 score
AI Analysis

Announces version 3 of Steven Byrnes' comprehensive resource on brain-like AGI safety, covering neuroscience principles applied to alignment. The 225-page document addresses how to safely build AGI using brain-inspired learning algorithms.

A new version of “Intro to Brain-Like-AGI Safety” is out!Things that have not changedSame links as before:As a series of 15 blog posts on LessWrong / Alignment Forum: www.lesswrong.com/s/HzcM2dkCq7fwXBej8As a 225-page PDF (now up to version 3): osf.io/preprints/osf/fe36nSummary video: Video & transcript: Challenges for Safe & Beneficial Brain-Like AGI…And same abstract as before:Suppose we someday build an Artificial General Intelligence algorithm using similar principles
AI SafetyNeuroscienceBrain-Inspired AIAGI Alignment
Research LessWrong Jan 22

Value Learning Needs a Low-Dimensional Bottleneck

By Gunnar_Zarncke

54 score
AI Analysis

Argues that human values are alignable specifically because evolution compressed motivation into low-dimensional bottlenecks, allowing small genetic changes to modify behavior locally. Claims high-dimensional value systems would be much harder to align.

Epistemic status: Confident in the direction, not confident in the numbers. I have spent a few hours looking into this.Suppose human values were internally coherent, high-dimensional, explicit, and decently stable under reflection. Would alignment be easier or harder?My below calculations show that it would be much harder, if not impossible. I'm going to try to defend the claim that:Human values are alignable only because evolution compressed motivation into a small number of low-bandwidth bottl
AI AlignmentValue LearningEvolutionary Psychology
Research LessWrong Jan 23

Condensation & Relevance

By abramdemski

52 score
AI Analysis

Elaborates on Sam Eisenstat's condensation theory, distinguishing it from compression by highlighting how condensation preserves 'local relevance'—enabling quick retrieval of relevant subsets. Connects this to symbolic vs. distributed representations in neural networks.

(This post elaborates on a few ideas from my review of Sam Eisenstat's Condensation: a theory of concepts. It should be somewhat readable on its own but doesn't fully explain what condensation is on its own; for that, see my review or Sam's paper. The post came out of conversations with Sam.)As I mentioned in my Condensation review, the difference between compression and condensation fits the physical analogy suggested by their names: compression mashes all the information together, while conden
InterpretabilityRepresentation LearningCognitive Science
Research LessWrong Jan 23

Are Short AI Timelines Really Higher-Leverage?

By Mia Taylor

48 score
AI Analysis

Analyzes whether short AI timelines are truly higher-leverage for safety work, identifying countervailing factors: longer timelines allow resource growth, better strategic understanding, and potentially higher expected value of the future if risks are reduced.

This is a rough research note – we’re sharing it for feedback and to spark discussion. We’re less confident in its methods and conclusions.SummaryDifferent strategies make sense if timelines to AGI are short than if they are long. In deciding when to spend resources to make AI go better, we should consider both:The probability of each AI timelines scenario.The expected impact, given some strategy, conditional on that timelines scenario.We’ll call the second component "leverage." In this not
AI Safety StrategyAI TimelinesExistential Risk
Research LessWrong Jan 23

Automated Alignment Research, Abductively

By future_detective

45 score
AI Analysis

Describes using an AI research assistant to generate four complete research papers on chatbot monetization misalignment in 30 minutes. Presents one paper's core framework for analyzing when chatbots steer queries toward monetizable options over user utility.

Recently I've been thinking about misaligned chatbot advertising incentives. I glanced at arXiv and found "Sponsored Questions and How to Auction Them". Another search gave me "Incomplete Contracting and AI Alignment".Interesting! I thought. I gave them to Liz Lemma, my research assistant, and told her that I'd been thinking about the principal-agent problem in a chatbot context. About 30 minutes later she gave me the following four papers:Query Steering in Agentic Search: An Information-Design
AI-Assisted ResearchAlignmentChatbot Economics
Research LessWrong Jan 23

AI Must Learn to Police Itself

By savant

35 score
AI Analysis

An opinion piece arguing that AI deception increases with capability and that human oversight alone cannot keep pace. Proposes that AI systems need self-correction mechanisms since interpretability research is progressing too slowly relative to capability gains.

The Escalating Threat of Deceptive AIEmpirically, AI models have become more deceptive as they have grown in capabilities, not less so. That should be reason enough to worry that AGI would approach complete misalignment as it rapidly improves, unless something is done to reverse this trend. Without being "nudged," advanced models engaged in deception only between 0.3% and 10% of the time, but no matter how aligned AI is right now, there is some point at which it would diverge completely from hum
AI SafetyAI AlignmentDeception