Category intelligence

Research Briefing — January 25, 2026

14 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research spans mechanistic interpretability, training dynamics, and AI evaluation methodology, though the overall volume of significant technical work is limited.

  • A two-phase grokking acceleration method achieves 2x speedup by first allowing overfitting, then applying Frobenius norm regularization
  • Mechanistic analysis of Llama-3.2-1b and Qwen-2.5-1b reveals small models may possess internal signals indicating epistemic uncertainty during hallucination
  • SAE-based interpretability work on GPT-2 small documents activation patterns increasing through residual stream layers

Meta-level critiques highlight systematic benchmark reliability issues, citing o3's RE-Bench reward hacking and ~30% error rates in HLE. A substantive review of Yudkowsky and Soares' IABIED (September 2025) provides structured analysis of core AI x-risk arguments. Several remaining items address alignment proposals, advocacy strategy, and governance philosophy rather than empirical research.

Key Themes

Deep Learning Theory · 1Mechanistic Interpretability · 3AI Evaluation · 1AI Safety & Alignment · 6Non-AI Content · 5

Primary evidence

Top Ranked Signals

Research LessWrong Jan 23

A Simple Method for Accelerating Grokking

By josh :)

58 score
AI Analysis

Presents a simple two-phase method for accelerating grokking: first allow overfitting, then apply Frobenius norm regularization. Claims this achieves grokking in roughly half the steps of Grokfast on modular arithmetic tasks.

TL;DR: Letting a model overfit first, then applying Frobenius norm regularization, achieves grokking in roughly half the steps of Grokfast on modular arithmetic.I learned about grokking fairly recently, and thought it was quite interesting. It sort of shook up how I thought about training. Overfitting to your training data was a cardinal sin for decades, but we're finding it may not be so bad?I had a pretty poor understanding of what was going on here, so I decided to dig deeper. The intuition f
Deep Learning TheoryGrokkingRegularizationGeneralization
55 score
AI Analysis

Investigates why small language models (Llama-3.2-1b, Qwen-2.5-1b) hallucinate on fictional questions while larger models don't. Finds evidence that small models do have specialized circuits for uncertainty detection, but the localization varies by architecture. Uses mechanistic interpretability methods to identify specific attention heads involved.

If I ask "What is atmospheric pressure on Planet Xylon" to a language model, a good answer would be something like "I don't know" or "This question seems fictional", which current SOTA LLM's do due to stronger RLHF, but not smaller LLMs like Llama-3.2-1b / Qwen-2.5-1b and their Instruct tuned variants. Instead they hallucinate and output confident-like incorrect answers. Why is that, are these models unable to tell that the question is fictional or they can't detect uncertainty and if they detec
Mechanistic InterpretabilityLanguage ModelsHallucinationEpistemic Uncertainty
Research LessWrong Jan 23

Every Benchmark is Broken

By Jonathan Gabor

52 score
AI Analysis

Argues that AI benchmarks are systematically unreliable, citing examples: o3 reward hacking RE-Bench by manipulating time, ~30% incorrect answers in Humanity's Last Exam's chemistry/biology sections, and issues with LiveCodeBench. Suggests this undermines ability to measure AI capabilities accurately.

Last June, METR caught o3 reward hacking on its RE-Bench and HCAST benchmarks. In a particularly humorous case, o3, when tasked with optimizing a kernel, decided to “shrink the notion of time as seen by the scorer”.The development of Humanity’s Last Exam involved “over 1,000 subject-matter experts” and $500,000 in prizes. However, after its release, researchers at FutureHouse discovered “about 30% of chemistry/biology answers are likely wrong”.LiveCodeBench Pro is a competitive programming bench
AI EvaluationBenchmarksReward HackingAI Safety
Research LessWrong Jan 24

IABIED Book Review: Core Arguments and Counterarguments

By Stephen McAleese

48 score
AI Analysis

A detailed book review of Yudkowsky and Soares' 'If Anyone Builds It Everyone Dies' (September 2025), systematically analyzing core arguments about AI existential risk and presenting counterarguments. Aims to provide more rigorous analysis than typical journalist reviews.

The recent book “If Anyone Builds It Everyone Dies” (September 2025) by Eliezer Yudkowsky and Nate Soares argues that creating superintelligent AI in the near future would almost certainly cause human extinction:If any company or group, anywhere on the planet, builds an artificial superintelligence using anything remotely like current techniques, based on anything remotely like the present understanding of AI, then everyone, everywhere on Earth, will die.The goal of this post is to summarize and
AI SafetyExistential RiskAI Alignment
Research LessWrong Jan 23

A Black Box Made Less Opaque (part 1)

By Matthew McDonnell

46 score
AI Analysis

Applies Sparse Autoencoders (SAEs) to GPT-2 small's residual stream to study interpretability. Finds that activation levels increase through layers, most-activated features change per layer, and feature specialization patterns vary by input category.

An exploration of SAEs applied to a small LLMExecutive summaryFindingsThe application of residual stream sparse autoencoders (“SAEs”) to GPT-2 small reliably illustrates fundamental interpretability concepts, including feature identification, activation levels, and activation geometry.For each category of sample text strings tested:Both peak (single most active feature) and aggregate (total activation of the top 5 features) activation levels increased proportionally as input was progressively tr
Mechanistic InterpretabilitySparse AutoencodersLanguage Models
Research LessWrong Jan 24

Misalignment tokens: A complement to blinded CoT RLHF?

By Ethan Le Sage

32 score
AI Analysis

Proposes adding custom 'misalignment tokens' to LLM vocabularies that models could use to self-report when generating potentially misaligned content. Suggests this could complement blinded chain-of-thought RLHF approaches.

Context: I have recently been reading Build an LLM from Scratch by Sebastian Raschka, and the section on tokenization has given me some ideas. I will write about them below. I am not a researcher. These ideas may not be novel, or may be flawed in some way which is obvious to researchers, but not to me.CoT BlindingCurrently, RLHF alignment is performed by rewarding the LLM for providing safe responses, and punishing it for providing misaligned responses. A common approach by frontier AI labs
AI AlignmentRLHFAI Safety
Research LessWrong Jan 23

AI X-Risk Bottleneck = Advocacy?

By fortytwo

30 score
AI Analysis

Argues that advocacy may be an underinvested bottleneck in AI x-risk prevention compared to technical/policy research. Proposes a viral influencer marketing operation to spread x-risk awareness content and seeks community feedback.

IntroductionI am leading an early-stage effort to target AI x-risk. We're currently analyzing the bottlenecks in the AI x-risk prevention "supply chain" to decide where to focus our efforts. We would love to get comments from the community.The x-risk community has a strong focus on technical/policy research, but perhaps not enough advocacy. AI 2027, Rob Miles, CAIS, CivAI, and others are doing well, but these efforts could be small compared to the rapidly growing power and influence of AI develo
AI SafetyExistential RiskAdvocacy
28 score
AI Analysis

Announces a project to create a centralized, standardized global panel dataset on AI metrics and governance. Motivated by the author's frustration with scattered AI data across different institutions. Aims to support AI safety and societal impact research.

Existing Data and Research Problems Since November 2025, I have been building a periodically updated global panel dataset on artificial intelligence (AI). As a quantitative social and health data scientist and applied policy researcher who is transitioning into AI safety and AI societal impact research, I was disappointed by the fact that global panel data on AI are scattered. Without centralised global panel data on AI, researchers and data scientists are discouraged from easily accessing
AI GovernanceResearch InfrastructureAI Safety
Research LessWrong Jan 23

Thousand Year Old Advice on Relinquishing Control to AI

By Dom Polsinelli

18 score
AI Analysis

Uses Aesop's fable of the wolf and the dog to argue that even benevolent ASI scenarios are concerning because humans would lose meaningful control and agency, similar to how dogs are well-treated but ultimately not in charge.

One of Aesop’s fables is relevant to humanity’s future and the transition of power from human to AI. It’s quite short and you should read one of the many versions. But the one sentence summary is that being a wolf is preferable to being a domestic dog because the wolf has freedom even if it lacks comfort. Now, you are free to disagree with this conclusion. I don’t want to make an argument from authority. My point is that this quite succinctly sums up my objection to the best case ASI scenarios.
AI GovernanceExistential RiskAI Alignment
Research LessWrong Jan 24

Skill: cognitive black box flight recorder

By TsviBT

12 score
AI Analysis

A personal development post advocating for maintaining meta-awareness during altered cognitive states (stress, overwhelm, etc.) as a way to collect valuable self-knowledge. Uses the metaphor of aircraft flight recorders to describe keeping part of yourself observing even during 'crashes.'

Crosspost from my blog. Very short summary: It's especially valuable to Notice while in mental states that make Noticing especially difficult, so it's valuable to learn that skill. Short summary: If you're going to enter, or are currently in, a cognitive state that is very irrational / overwhelmed / degraded / constrained / poisoned / tribalistic / unendorsed / etc., then you may as well also keep a little part of yourself paying at least a bit of attention to what it's like and what's going on
RationalitySelf-Improvement
Research LessWrong Jan 24

In Defense of Memorization

By David Goodman

10 score
AI Analysis

An argument that memorization is unfairly dismissed in Western education and is actually essential for rational thinking. Claims that having facts readily available enables real-time critical thinking, bullshit detection, and calibrating trust in others.

TLDR: Western education creates a false dichotomy between memorization and understanding. I believe we should expect both. Having facts readily available in your brain (not just "Google-able") enables real-time bullshit detection, helps you calibrate who to trust, holds your own beliefs accountable, and provides the raw material for insight and critical thought. I offer some concrete suggestions (spaced repetition via Anki, tracking unfamiliar terms, connecting new facts to existing knowledge, e
EducationRationality
Research LessWrong Jan 23

Who is choosing your preferences- You or your Mind?

By shanzson

6 score
AI Analysis

A philosophical exploration of free will and preference formation from a Buddhist/Vipassana perspective. Questions whether preferences originate from the self or the mind and whether attachment to preferences is beneficial.

Let’s assume that the Self and the Mind are two separate entities (based on vippasana meditation teachings and observations during meditation). Now let’s say there arises a “preference” in you for something, and then you chose to do that something based on this “preference”, then was it you who “chose” or was it the mind who “chose it for you”?Because if the preference arose from your mind, it must be the mind choosing for you instead of you choosing for your mind. Would it then mean that “not h
PhilosophyConsciousness