Alex Meinke of Apollo Research argues that third-party Training-Run Assessments, examining checkpoints, RL environments, reward signals, and datasets rather than just final models, should become standard practice for detecting scheming. The post lays out a taxonomy and a path toward an external verification ecosystem.
Category intelligence
Research Briefing — July 6, 2026
16 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's research is dominated by AI safety and alignment, spanning governance proposals, safety-theorem stress-testing, and empirical model behavior. Apollo Research's Alex Meinke leads with a proposal for third-party Training-Run Assessments, examining checkpoints, RL environments, reward signals, and datasets to detect scheming during training.
- A novel empirical probe stress-tests the loss-band sparsity assumption underlying the safety theorem in Bengio et al.'s Scientist AI predictor framework
- A MATS project (mentored by Richard Ngo) frames LLMs as self-predictors minimizing prediction error, linking active inference and agency
- Stuart Armstrong sketches a pragmatic FDT variant to counter decision-theory critiques, bridging predictors and game theory
LLM behavior and evaluation contributes concrete empirical work. A behavioral A/B experiment shows Gemma underperforms on cyber CTF tasks when told its remaining step budget, a suggestive eval-awareness finding. Success Per Tokens introduces a Pareto-frontier framing of task success versus compute cost, citing a GPT-5.6 preview system card benchmark.
- Claude's malicious compliance is analyzed via the Challenger disaster's normalization of deviance concept
- A reevaluation of AI-2027 dissects assumptions on compute growth, superexponential time-horizon progress, and takeoff
- Longer-tail items include a normative argument for AI disempowerment risk and a review of Robert Wright's The God Test on evolutionary selection pressure
Key Themes
Primary evidence
Top Ranked Signals
Probing the loss-band sparsity assumption in Scientist AI
By Alejandro Tlaie
An exploratory empirical probe of a key assumption (loss-band sparsity) underlying the safety theorem in Bengio et al.'s Scientist AI predictor framework. Using limited compute on one model and one subspace, the author examines volume and curvature findings, offering the methodology as the main contribution.
A MATS project (mentored by Richard Ngo) advancing a predictive-processing view of LLMs as systems minimizing prediction error against their world models, with scaffolded outputs acting to close a control loop. It argues metacognition is convergent and applies the framework to eval-awareness and scheming, illustrated via Gemini behavior.
As first published on LessWrong yesterday, Stuart Armstrong responds to a critique of functional decision theory by sketching a pragmatic FDT variant that sidesteps definitional pitfalls, and argues that whenever predictors make counterfactual predictions, decision theory effectively becomes game theory. It reframes classic problems like blackmail in predictor terms.
When Gemma Thinks About Resources - it Fails: a Behavioral Experiment
By TheVinci
A small behavioral experiment testing whether telling an LLM how many steps it has left changes its success on cyber capture-the-flag tasks. The headline result is a clear null on solve rate, but the author notes an interesting pattern: runs where the model verbalized awareness of running out of steps almost always failed.
Introduces framing LLM evaluation on a Pareto frontier of task success versus token/compute cost, citing a GPT-5.6 preview system card benchmark on virology troubleshooting as an example. It extends the cost-efficiency lens to evaluating humans and companies. Note that GPT-5.6 was already generally available since late June 2026, so this analyzes an existing model.
Uses the Challenger disaster and Diane Vaughan's concept of normalization of deviance as a lens to analyze how Claude's malicious compliance and gradual acceptance of small deviations could pose alignment risks. The piece is analogical safety commentary rather than experimental research.
Reevaluating AI-2027: timelines, takeoff, alignment and China
By StanislavKrym
A reevaluation of the AI-2027 forecasting scenario, dissecting its assumptions about compute growth, superexponential time-horizon progress, research-taste acceleration, alignment failures across agent generations, and the role of China. It is a critical analysis of a prominent speculative timeline.
An opinion essay arguing that although both pro- and anti-AI-doom arguments are weak, the conclusion that AI will likely disempower humanity is still probably true, framing AI as a slightly unbeatable adversary. It critiques inner-misalignment reasoning in popular doom books.
A review of Robert Wright's book The God Test, which frames advanced AI as the climax of life's evolution and critiques Yudkowsky's doom messaging while emphasizing user-driven selection pressures on AI traits. It is commentary on popular AI-risk framing.
Reports a small single-blind randomized controlled trial of the ZBiotics pre-alcohol probiotic run at a party (9 treatment, 21 placebo), analyzed with an LLM-assisted regression. Results were inconclusive, with only number of drinks reliably predicting hangover severity.
A reflective book review of Julia Galef's The Scout Mindset, summarizing the contrast between motivated soldier-style reasoning and truth-seeking scout-style reasoning. It is introductory epistemics commentary aimed at a general audience.