Presents a simple two-phase method for accelerating grokking: first allow overfitting, then apply Frobenius norm regularization. Claims this achieves grokking in roughly half the steps of Grokfast on modular arithmetic tasks.
Category intelligence
Research Briefing — January 25, 2026
14 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's research spans mechanistic interpretability, training dynamics, and AI evaluation methodology, though the overall volume of significant technical work is limited.
- A two-phase grokking acceleration method achieves 2x speedup by first allowing overfitting, then applying Frobenius norm regularization
- Mechanistic analysis of Llama-3.2-1b and Qwen-2.5-1b reveals small models may possess internal signals indicating epistemic uncertainty during hallucination
- SAE-based interpretability work on GPT-2 small documents activation patterns increasing through residual stream layers
Meta-level critiques highlight systematic benchmark reliability issues, citing o3's RE-Bench reward hacking and ~30% error rates in HLE. A substantive review of Yudkowsky and Soares' IABIED (September 2025) provides structured analysis of core AI x-risk arguments. Several remaining items address alignment proposals, advocacy strategy, and governance philosophy rather than empirical research.
Key Themes
Primary evidence
Top Ranked Signals
Investigates why small language models (Llama-3.2-1b, Qwen-2.5-1b) hallucinate on fictional questions while larger models don't. Finds evidence that small models do have specialized circuits for uncertainty detection, but the localization varies by architecture. Uses mechanistic interpretability methods to identify specific attention heads involved.
Argues that AI benchmarks are systematically unreliable, citing examples: o3 reward hacking RE-Bench by manipulating time, ~30% incorrect answers in Humanity's Last Exam's chemistry/biology sections, and issues with LiveCodeBench. Suggests this undermines ability to measure AI capabilities accurately.
IABIED Book Review: Core Arguments and Counterarguments
By Stephen McAleese
A detailed book review of Yudkowsky and Soares' 'If Anyone Builds It Everyone Dies' (September 2025), systematically analyzing core arguments about AI existential risk and presenting counterarguments. Aims to provide more rigorous analysis than typical journalist reviews.
Applies Sparse Autoencoders (SAEs) to GPT-2 small's residual stream to study interpretability. Finds that activation levels increase through layers, most-activated features change per layer, and feature specialization patterns vary by input category.
Proposes adding custom 'misalignment tokens' to LLM vocabularies that models could use to self-report when generating potentially misaligned content. Suggests this could complement blinded chain-of-thought RLHF approaches.
Argues that advocacy may be an underinvested bottleneck in AI x-risk prevention compared to technical/policy research. Proposes a viral influencer marketing operation to spread x-risk awareness content and seeks community feedback.
The Global AI Dataset (GAID) Project: From Closing Research Gaps to Building Responsible and Trustworthy AI
By Jason Hung
Announces a project to create a centralized, standardized global panel dataset on AI metrics and governance. Motivated by the author's frustration with scattered AI data across different institutions. Aims to support AI safety and societal impact research.
Uses Aesop's fable of the wolf and the dog to argue that even benevolent ASI scenarios are concerning because humans would lose meaningful control and agency, similar to how dogs are well-treated but ultimately not in charge.
A personal development post advocating for maintaining meta-awareness during altered cognitive states (stress, overwhelm, etc.) as a way to collect valuable self-knowledge. Uses the metaphor of aircraft flight recorders to describe keeping part of yourself observing even during 'crashes.'
An argument that memorization is unfairly dismissed in Western education and is actually essential for rational thinking. Claims that having facts readily available enables real-time critical thinking, bullshit detection, and calibrating trust in others.
A philosophical exploration of free will and preference formation from a Buddhist/Vipassana perspective. Questions whether preferences originate from the self or the mind and whether attachment to preferences is beneficial.