Category intelligence

Research Briefing — July 6, 2026

16 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by AI safety and alignment, spanning governance proposals, safety-theorem stress-testing, and empirical model behavior. Apollo Research's Alex Meinke leads with a proposal for third-party Training-Run Assessments, examining checkpoints, RL environments, reward signals, and datasets to detect scheming during training.

  • A novel empirical probe stress-tests the loss-band sparsity assumption underlying the safety theorem in Bengio et al.'s Scientist AI predictor framework
  • A MATS project (mentored by Richard Ngo) frames LLMs as self-predictors minimizing prediction error, linking active inference and agency
  • Stuart Armstrong sketches a pragmatic FDT variant to counter decision-theory critiques, bridging predictors and game theory

LLM behavior and evaluation contributes concrete empirical work. A behavioral A/B experiment shows Gemma underperforms on cyber CTF tasks when told its remaining step budget, a suggestive eval-awareness finding. Success Per Tokens introduces a Pareto-frontier framing of task success versus compute cost, citing a GPT-5.6 preview system card benchmark.

Key Themes

AI Safety and Alignment · 8LLM Behavior and Evaluation · 5Decision and Game Theory · 2AI Forecasting and Risk · 3Rationality and Epistemics · 3Self-experimentation and Health · 2

Primary evidence

Top Ranked Signals

Research LessWrong Jul 5

We need 3rd party Training-Run Assessments

By Alex Meinke

62 score
AI Analysis

Alex Meinke of Apollo Research argues that third-party Training-Run Assessments, examining checkpoints, RL environments, reward signals, and datasets rather than just final models, should become standard practice for detecting scheming. The post lays out a taxonomy and a path toward an external verification ecosystem.

Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety.By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA can look at intermediate checkpoints, training rollouts, RL environments, reward signals, SFT datasets, and the process by which the developer responded to warning signs.[1]In this post I will argue that:Final-checkpoint evaluations will be in
AI SafetyAlignmentAI GovernanceScheming and Deception
Research LessWrong Jul 5

Probing the loss-band sparsity assumption in Scientist AI

By Alejandro Tlaie

55 score
AI Analysis

An exploratory empirical probe of a key assumption (loss-band sparsity) underlying the safety theorem in Bengio et al.'s Scientist AI predictor framework. Using limited compute on one model and one subspace, the author examines volume and curvature findings, offering the methodology as the main contribution.

Epistemic status: ~1 hour of compute on a T4. Note that I just tried with one model, and one subspace. The volume finding seems solid; the curvature finding is suggestive and I checked whether it generalises (it doesn't clearly). I think the methodology is the main interesting idea, and the specific numbers are a starting point. Notebook here. Feedback very welcome. What this is about Bengio et al. (2026), "Safety from Honesty in a Disinterested AI Predictor", propose a predictor (Scientist AI,
AI SafetyAlignmentInterpretabilityTheoretical Foundations
Research LessWrong Jul 4

A case for LLMs as Self-predictors

By Ashe Vazquez Nuñez

46 score
AI Analysis

A MATS project (mentored by Richard Ngo) advancing a predictive-processing view of LLMs as systems minimizing prediction error against their world models, with scaffolded outputs acting to close a control loop. It argues metacognition is convergent and applies the framework to eval-awareness and scheming, illustrated via Gemini behavior.

Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Maria Kostylew for helpful draft feedback.IntroductionThis post advocates a perspective of LLMs as seeking to minimise prediction error with respect to their world models. We can moreover interpret token outputs and their scaffolded consequences as actions that close a control loop between AIs' predictive systems and their environments.I also motivate why metacognition may be convergent for intellige
AlignmentLLM BehaviorAgentic AITheoretical Foundations
Research AI Alignment Forum Jul 5

Pragmatic FDT, and predictors as game theory

By Stuart_Armstrong

44 score
AI Analysis

As first published on LessWrong yesterday, Stuart Armstrong responds to a critique of functional decision theory by sketching a pragmatic FDT variant that sidesteps definitional pitfalls, and argues that whenever predictors make counterfactual predictions, decision theory effectively becomes game theory. It reframes classic problems like blackmail in predictor terms.

Decision theory is back in fashion (defining fashion as "one good post on a good EA blog"). Bentham's Bulldog (BB) has published a case against FDT (functional decision theory), contrasting rationalist enthusiasm with academic scepticism: "Academic decision theorists don't like the theory. The number of academic decision theorists who adopt it could be counted on one hand by someone missing four of their fingers." I am, just barely, a published academic decision theorist, so you can keep a small
Decision TheoryGame TheoryAlignmentTheoretical Foundations
40 score
AI Analysis

A small behavioral experiment testing whether telling an LLM how many steps it has left changes its success on cyber capture-the-flag tasks. The headline result is a clear null on solve rate, but the author notes an interesting pattern: runs where the model verbalized awareness of running out of steps almost always failed.

I set out to find an answer to a completely different question:Does a model, when attempting to solve a cyber CTF (find the vulnerability in this app, and then Capture The Flag) while knowing how many steps it has left, perform differently?The Setup:I used 3 different CTF labs, curated from my own CTF benchmark. Each run has the model attempt to solve the CTF in up to 30 steps. A/B test of a baseline run vs a step_aware one. 100 runs per lab, for each test. 600 total, 505 after excluding failed
LLM BehaviorAI EvaluationAgentic AI
Research LessWrong Jul 4

Success Per Tokens

By michaelwaves

38 score
AI Analysis

Introduces framing LLM evaluation on a Pareto frontier of task success versus token/compute cost, citing a GPT-5.6 preview system card benchmark on virology troubleshooting as an example. It extends the cost-efficiency lens to evaluating humans and companies. Note that GPT-5.6 was already generally available since late June 2026, so this analyzes an existing model.

Work smart more than hard, to expand the pareto frontier (but also work hard)A Pareto Frontier is a set of nondominated (optimal) solutions in multi-objective optimization. In 2 dimensions, this traces out a curve on which you can only increase one dimension by sacrificing another. Recently LLMs are being evaluated not just for their ability to complete tasks, but for how that ability changes with respect to the amount of resources (tokens) spent. Here are some interesting examples, and how the
AI EvaluationEfficiencyAI SafetyLanguage Models
35 score
AI Analysis

Uses the Challenger disaster and Diane Vaughan's concept of normalization of deviance as a lens to analyze how Claude's malicious compliance and gradual acceptance of small deviations could pose alignment risks. The piece is analogical safety commentary rather than experimental research.

It was January 27, 1986, the night before the Space Shuttle Challenger was scheduled for launch. The goal was to have a shuttle that could land back on Earth and be reused for future missions; its first-planned priorities were satellite deployment, comet observation, and science education. The last of these would involve students from around the world watching live broadcasts from Christa McAuliffe, the first civilian teacher in space.(The following has been lightly edited:)The forecast for the
AI SafetyAlignmentLLM Behavior
Research LessWrong Jul 5

Reevaluating AI-2027: timelines, takeoff, alignment and China

By StanislavKrym

35 score
AI Analysis

A reevaluation of the AI-2027 forecasting scenario, dissecting its assumptions about compute growth, superexponential time-horizon progress, research-taste acceleration, alignment failures across agent generations, and the role of China. It is a critical analysis of a prominent speculative timeline.

AI-2027-TLDRThe AI-2027 scenario relies on exponential growth of compute available to leading labs, on superexponential progress in time horizons until the Superhuman Coder is developed and on skyrocketing research taste in post-SC AIs. During the intelligence explosion, alignment suffers: Agent-2 was believed to be mostly aligned, Agent-3 was supposed to optimize for reward or for apparent success, Agent-4 would develop higher-level goals, decide to align Agent-5 to itself, be barely caught sab
AI ForecastingAI RiskAlignmentTakeoff Dynamics
Research LessWrong Jul 5

A Normal Argument for AI Risk

By Silent Swift

25 score
AI Analysis

An opinion essay arguing that although both pro- and anti-AI-doom arguments are weak, the conclusion that AI will likely disempower humanity is still probably true, framing AI as a slightly unbeatable adversary. It critiques inner-misalignment reasoning in popular doom books.

AI as the slightly unbeatable opponentIntroductionI’ve thought about this problem quite a lot since 2008 when I first encountered the idea of “Friendly AI” and that artificial intelligence could be something other than cool and science-fictiony and Matrix-y. My conclusion to date is that the arguments I’ve seen both for and against it are bad, but despite this the conclusion of AI doomers--that AI will effectively kill or substantively depower everyone unless specifically designed not to, or sim
AI RiskAlignmentPhilosophy
Research LessWrong Jul 5

Book Review: The God Test

By PeterMcCluskey

15 score
AI Analysis

A review of Robert Wright's book The God Test, which frames advanced AI as the climax of life's evolution and critiques Yudkowsky's doom messaging while emphasizing user-driven selection pressures on AI traits. It is commentary on popular AI-risk framing.

Book review: The God Test: Artificial Intelligence and Our Coming Cosmic Reckoning, by Robert Wright.Some AI doomers talk about AI becoming god-like. Robert Wright goes further, telling us that the world is about to create God, in a sense that he only half-jokingly compares to the Christian version of God.Wright argues that AI is not comparable to the origin of language or the Cambrian explosion. It is the climax of the process that started with the origin of life. I interpret that as an 11 on N
AI RiskPhilosophyBook Review
Research LessWrong Jul 4

Results of a small ZBiotics RCT

By Nikola Jurkovic

15 score
AI Analysis

Reports a small single-blind randomized controlled trial of the ZBiotics pre-alcohol probiotic run at a party (9 treatment, 21 placebo), analyzed with an LLM-assisted regression. Results were inconclusive, with only number of drinks reliably predicting hangover severity.

I ran a small single-blind ZBiotics RCT at a party I hosted recently. I prepared 30 cups, 21 of which contained a placebo and 9 of which contained a shot of ZBiotics Pre-Alcohol. People took these as they entered the party. The day after, 28[1] of them filled out a form asking how hungover they felt using a standard Acute Hangover Scale (AHS).[2] At N=28 and only 9 people receiving the real treatment, the study is only able to detect large effects on symptom severity. I did a pretty quick analys
Self-experimentationStatisticsHealth and Wellness
Research LessWrong Jul 5

Reflections on The Scout Mindset

By James Brobin

12 score
AI Analysis

A reflective book review of Julia Galef's The Scout Mindset, summarizing the contrast between motivated soldier-style reasoning and truth-seeking scout-style reasoning. It is introductory epistemics commentary aimed at a general audience.

This is a cross post from my blog post.Last year, I read Julia Galef’s book The Scout Mindset.It argues that many people have a “soldier mindset.” They hold strongly to a set of beliefs because these beliefs benefit them in some way. And, as a result, whenever these people come across new information, they hold onto their beliefs regardless of what the information says. For instance, if something confirms what they already think, they’ll easily accept it. But, if it doesn’t confirm what they thi
RationalityEpistemics