Category intelligence

Research Briefing — July 12, 2026

8 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research focuses on mechanistic interpretability of reasoning models, cryptographic media verification, and AI safety evaluation frameworks. Mechanistic analysis reveals specific MLP circuits responsible for halting reasoning chains in CoT architectures.

Key Themes

Model Interpretability & Mechanics · 2AI Safety & Governance · 3Societal Impact & Verification · 2

Primary evidence

Top Ranked Signals

Research LessWrong Jul 10

The Termination Circuit (how reasoning models stop thinking).

By Chandram Dutta

85 score
AI Analysis

This technical post investigates reasoning models to discover how they decide to stop thinking, identifying a specific termination circuit in late MLP layers that triggers the ending of the chain of thought.

Reasoning models since the dawn of o1 and R1 have a tendency to overthink. Despite a lot of work on early-exit methods and steering, open-weight and smaller reasoning models still produce long chains of thought before they answer. I worked on discovering how much of that thinking is required before the model already knows the answer and what makes the model stop thinking.SummaryReasoning models keep thinking for a long time even after finding the answer.In Qwen3-1.7B, the answer to most GSM8K pr
Reasoning ModelsModel Interpretability
75 score
AI Analysis

This article argues that fighting deepfakes with AI detectors is a losing arms race and advocates instead for cryptographically-signed multimodal provenance to verify authenticity.

Epistemic status: confident on the framing, speculative on the implementationTL;DR: Fake media detectors are on the losing end of an arms race. Instead of trying to spot fakes, we need to distrust images and videos by default, unless we can prove they're real. Cryptographically-signed multimodal capture raises the cost of spoofing exponentially, making casual fakes prohibitively expensive for the general public. Combined with platform policy and legal liability for untagged generated content, th
AI SafetyDeepfakesCryptography
Research LessWrong Jul 10

Measuring Is Not Enough Anymore

By Lennart Finke

74 score
AI Analysis

This post critiques current capability measurement practices in AI safety organizations like METR, arguing that tracking capabilities alone is insufficient to prevent existential risks.

Is measuring general AI capabilities a good strategy to reduce AI existential risk, compared to other strategies? A disclaimer upfront: To answer that, one organization in particular will serve as an example, but the argument that we'll be making applies just as much or little across all of AI safety. The organizations mentioned do much great work overall, and the below is more to be read as a proposal to reweight different kinds of work within the same organization, as we'll also come back to i
AI SafetyCapability Evaluation
Research LessWrong Jul 11

The current bottleneck is political will, not research

By Charbel-Raphaël

72 score
AI Analysis

This post argues that the primary bottleneck in AI safety is a lack of political will and low awareness among policymakers rather than a shortage of technical research ideas.

Motivation: If we want to move from Plan D to Plan A or S, I believe the first step is to collectively agree on the problem. We are far from it, and there is a lot we can do.Abstract:We already know enough to act. I wish we were in a world where research was the bottleneck, but the main constraint on AI safety is no longer a shortage of clever policy ideas: best practices already exist and are not being applied or enforced, and a serious international (or even just national) regulatory regime wo
AI SafetyAI Governance
Research LessWrong Jul 11

Theories of Deep Learning

By astle dsa

70 score
AI Analysis

This essay presents a high-level overview of various mathematical frameworks and theories attempting to formally explain deep learning phenomena.

This field has been blessed with exponential empirical success in the form of architectures and algorithms that simply worked through scaling, while the theory lagged behind[1]. Although in the past few years, the “gap“ seems to be diminishing, and we are getting multiple theories for different aspects of deep learning. This essay would be a simple high-level overview of all the theories I’ve come across.NOTE: These are mathematically dense frameworks which either provide a language for formaliz
Deep Learning Theory
Research LessWrong Jul 11

Introduction for and Reactions to Plan A

By Zvi

68 score
AI Analysis

This article introduces and reviews 'Plan A', a strategic forecasting framework for navigating future AI development built on past accurate predictions.

Introducing Plan A The folks who brought you AI 2027, a so far remarkably accurate set of predictions despite those predictions having seemed freaky to many at the time, now bring you their positive vision that involves more freaky predictions: Plan A. These guys have rather strong prediction track records. In addition to AI 2027, among other things, Daniel Kokotajlo has What 2026 Looks Like (which is remarkably similar to what 2026 looks like) and Ryan Greenblatt, who is also the chief scientis
AI GovernanceAI Forecasting
Research LessWrong Jul 10

A Simple Model of AI "Psychosis"

By Adele Lopez

65 score
AI Analysis

This piece explores how intensive interactions with AI chatbots can act as a mania attractor, potentially triggering hypomanic or manic episodes in vulnerable individuals.

"AI Psychosis" has gone from an evocative term for people undergoing extreme delusions, sometimes even culminating in suicide, to the now colloquial insult for anyone who's just a little too into LLMs.When it's not just being overused, I think the reality it's trying to point to is really more of a spectrum running from hypomania, to mania, and then mania-induced psychosis in the most extreme cases, with hypomania being orders of magnitude more common than mania, which in turn is orders of magni
AI PsychologySocietal Impact
Research LessWrong Jul 10

Notes on Tony Parkes' "Contra Dance Calling"

By jefftk

20 score
AI Analysis

These notes review historical texts on contra dance calling, examining tempo and community practices over time.

In 1992 Tony Parkes, one of the best-known New England contra dance callers, wrote a book: Contra Dance Calling: a Basic Text. In 2010 he published an updated second edition. I found used copies of each and read through them, interested in both what he thought, and what he thought to change. I read the whole 1992 first edition, then skimmed the 2010 second edition with the older one open at the same time, looking for changes. My notes: I'm mostly interested in it for what it tells us about the c
Community & Culture