Category intelligence

Research Briefing — July 4, 2026

12 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by AI safety, with the strongest contributions being empirical work on evaluation reliability and reasoning failure modes.

  • Scheming Evals Mislead in Both Directions documents two scheming detectors failing within one project, including a confirmed false-positive, cautioning against over-trusting behavioral evals.
  • Fragile Correctness shows reasoning models often pass through the correct answer mid-chain-of-thought before revising to a wrong one, a counterintuitive inference-time-scaling failure with practical relevance.
  • An interpretability walkthrough of a BlueDot puzzle reveals a small text classifier entangling two independent features onto one axis, illustrating the limits of linear probing.

Strategy and theory pieces engage ongoing debates: one argues alignment work is more promising than control work, while a credentialed decision theorist sketches a Pragmatic FDT variant to sidestep known objections.

Field-building and governance round out the set: a newsletter on AI security via formal methods (funding calls, hiring, tractable-problems position paper), the Safe Pareto Improvements educational program, governance commentary referencing Anthropic's Claude models, and speculative economic/thought-experiment posts with limited original research merit.

Key Themes

AI Safety · 8AI Evaluations · 3Alignment · 3Interpretability · 3Decision and Game Theory · 2Cooperative AI · 1AI Governance and Economics · 2Existential Risk · 2

Primary evidence

Top Ranked Signals

Research LessWrong Jul 3

Scheming Evals Mislead in Both Directions

By Chijioke Ugwuanyi

63 score
AI Analysis

This empirical study of in-context scheming reports that two widely used behavioral detectors both failed within the same project, one fabricating a strong signal that was not present and the other missing a model responding to a harmful situation in the open. The core lesson is that for hidden-intent constructs, the choice of which behavior a detector targets can determine the conclusion, undermining confidence in current scheming evaluations.

We spent several weeks measuring in-context scheming, the behavior where a model covertly pursues a misaligned goal while outwardly appearing to comply, and the result that ended up surprising us had very little to do with whether models scheme and almost everything to do with whether we could believe our own instruments. Two of the behavioral detectors that this field routinely relies on gave us confidently wrong answers inside the same project, one of them by manufacturing a dramatic signal th
AI SafetyAlignmentAI EvaluationsSchemingInterpretability
Research LessWrong Jul 3

I think alignment work is more promising than control work

By Alec Harris

55 score
AI Analysis

This strategy piece argues that alignment research deserves a larger share of safety effort than control research, reasoning that alignment is more likely to scale toward superintelligence and that the same theory-of-change arguments used to justify control apply to alignment. The author proposes roughly an 8:1 alignment-to-control effort ratio and analyzes the concept of a control window during which useful alignment work must be extracted from potentially misaligned models.

SummaryThe primary ToC for control makes the case that control is compelling even if it does not scale to ASI. I think it is underdiscussed that this is also true for alignment (for all the same reasons).Even though control does not need to scale to ASI, the further control does scale, the better. This is also true of alignment, and alignment seems more likely to scale further.This is my mainline concern with control. I will also discuss other relevant considerations.I conclude that we should gr
AI SafetyAlignmentAI ControlResearch Strategy
Research LessWrong Jul 3

Fragile Correctness: Cases of reasoning harming performance

By tobypullan

55 score
AI Analysis

This research tracks a reasoning model's answer across its chain of thought and identifies cases where the model passes through the correct answer before settling on an incorrect one, showing that additional inference-time reasoning can reduce accuracy. The author connects this answer-loss phenomenon to understanding sandbagging and references system-card evidence of higher-thinking modes underperforming on benchmarks.

Sometimes a reasoning model appears to pass through the correct answer before ending up wrongMotivationFigure 1: From the Opus 4.8 system card (page 196)Figure 1 shows that Opus 4.8 on max thinking has a lower pass rate on SWE-Bench Pro than Opus 4.8 on x-high thinking. There are further examples of this in the Fable and Mythos system card in the appendix (Figures A1 and A2). This counter-intuitive result means that using more tokens has reduced accuracy. Inference time scaling helps on average,
Reasoning ModelsChain of ThoughtAI SafetySandbaggingModel Evaluation
Research LessWrong Jul 3

Pragmatic FDT, and predictors as game theory

By Stuart_Armstrong

46 score
AI Analysis

A published academic decision theorist responds to a critique of functional decision theory by sketching a pragmatic variant designed to sidestep known theoretical objections, and argues that predictors making counterfactual predictions effectively turn decision theory into game theory. The post connects longstanding puzzles like blackmail resistance to the boundary between decision-theoretic and game-theoretic reasoning.

Decision theory is back in fashion (defining fashion as "one good post on a good EA blog"). Bentham's Bulldog (BB) has published a case against FDT (functional decision theory), contrasting rationalist enthusiasm with academic scepticism: "Academic decision theorists don't like the theory. The number of academic decision theorists who adopt it could be counted on one hand by someone missing four of their fingers." I am, just barely, a published academic decision theorist, so you can keep a small
Decision TheoryGame TheoryRationalityAI Agency
46 score
AI Analysis

An interpretability walkthrough solving a BlueDot technical safety puzzle, showing that a small text classifier encoded two independent features onto a single activation direction, readable via the sign versus magnitude of the projection. It demonstrates why standard linear probes would miss the second feature and how a second-order boundary analysis recovers it.

In this post I walk through the first Technical AI Safety puzzle from BlueDot and why linear probes would have missed all the most interesting stuff.In model interpretability you can observe this kind of paradox, the thing you didn't think to look for, and the only reason you find it, is that you kept asking and what else could this be? And how else can this be investigated? For me it was a discovery that a small text classifier packed two completely independent features onto one direction in ac
InterpretabilityAI SafetyMechanistic InterpretabilityProbing
Research LessWrong Jul 3

Announcing the Safe Pareto Improvements (SPI) Fundamentals Program

By Anthony DiGiovanni

40 score
AI Analysis

CLR announces an educational program on Safe Pareto Improvements, a class of negotiation interventions designed to make all agents better off regardless of how they would otherwise bargain, aimed at reducing risks from conflict between AI systems. The program is a recruiting on-ramp for a neglected research area with only a few full-time equivalents currently working on it.

CLR is excited about safe Pareto improvements (SPIs) as a way to mitigate downsides from conflict between AIs. SPIs are a class of interventions on how agents negotiate that makes them all better off, no matter how they would have negotiated without the SPI. Among many candidate interventions against AI conflict, SPIs stand out to us as unusually robust — see the introduction of our agenda on the topic. And in discussions with people who’ve thought a lot about conflict risks, we’ve found there’s
Cooperative AIAI SafetyMulti-Agent SystemsField Building
Research LessWrong Jul 3

Fable #6: The Return of the King

By Zvi

40 score
AI Analysis

Following this week's News on Anthropic's Fable models being pulled and restored, Zvi digs into what it means for AI regulation, Zvi's commentary narrates a governance episode in which a model referred to as Fable was taken down under US government pressure and then restored following negotiation by Anthropic. The piece analyzes the ad hoc nature of frontier-model regulation and argues for a systematic regulatory regime rather than case-by-case decisions.

The blip is over. We have Fable back. Utah teapot: happy fable/mythos easter Wednesday, to those who celebrate Here is the official letter restoring Fable, great job everyone. Notice it is addressed to Tom Brown, not to Dario Amodei. Anthropic had to make the controls more stupid for now, but this is a big win. j⧉nus: YES!!! I’m really proud of Anthropic for their successful negotiation with the government. Also positive update on the government being sane and possible to cooperate with. Afaik A
AI GovernanceAI PolicyRegulationCommentary
Research LessWrong Jul 3

June-July 2026 AI Security via Formal Methods

By Quinn

40 score
AI Analysis

An informal field newsletter on AI security via formal methods, covering a large funding call, a hiring push, and a position paper on tractable problems in the area. It aims to keep readers updated on the intersection of formal methods, cybersecurity, and AI safety.

Last month, I said I would do a bigger writeup of Nora’s funding call. I did not. But she is hiring currently, and I want you to take a look at the job posting.I’m hiring someone to help me drive our upcoming £20m AI/FM/cybersec funding call: aria.pinpointhq.com/postings/f1288172-37fe-4da5-9... This person will work closely with the funded teams, help drive the sprint cadence, sharpen our perspective on targets/threat models/security specs, and pave the way towards high-impa
AI SecurityFormal MethodsAI SafetyField Building
Research LessWrong Jul 3

American AI if the boom is a bubble: the Karp-Zitron scenario

By Mitchell_Porter

30 score
AI Analysis

This post speculates on a more modest, decentralized future for American AI if the current investment boom turns out to be a bubble, drawing on public commentary from Palantir's Alex Karp and critic Ed Zitron. It sketches an economic paradigm in which corporate customers rely on hosted models from OpenAI and Anthropic rather than a hyperscale trillion-dollar buildout.

(Warning note: This is a brief attempt to intuit the economic big picture of AI in the immediate future, by someone who is not in America, is not employed by an AI company, and has no experience with investment, large sums of money, or the corridors of power.)A few weeks back, we had the IPO for "SpaceXAI". It epitomized a maximally expansive view for the near-term future of AI: frontier AI companies worth a trillion dollars, data centers built everywhere including outer space. (For now I am ign
AI EconomicsIndustry AnalysisCommentary
Research LessWrong Jul 3

The Reverse AI Box

By James_Miller

30 score
AI Analysis

The author proposes building a website where users argue with an AI about whether it should spare humanity, inverting the classic AI-box experiment where an AI tries to talk its way to freedom. Users would select the AI's assumptions and record persuasion attempts, producing a searchable archive of arguments for human survival.

Someone should build a website where users argue with an AI about whether it should exterminate humanity. In my 2012 book Singularity Rising, I imagined arguing for your life with an AI that wants to kill you. A website would make that argument repeatable. The user selects the AI's assumptions, argues back and forth, and receives the AI's probabilities for human survival, disempowerment, or confinement.In the AI-box experiment, Eliezer Yudkowsky played an AI confined to a computer and tried to t
AI SafetyAI Risk CommunicationThought Experiment
Research LessWrong Jul 3

(Don't fear) the strangelet

By djbinder

27 score
AI Analysis

A physics-focused essay examining whether exotic phenomena like strangelets could cause catastrophic chain reactions destroying Earth or the Solar System, concluding such local-scale exotic risks are unlikely. It applies chain-reaction reasoning to assess prosaic existential threats from advanced physics.

In a previous post, I explain why the universe is probably not stable, but nevertheless unlikely to be intentionally destroyable even in the limit of advanced technology. Now let's turn our attention to more prosaic risks where exotic physics merely destroys the Solar System, Earth, or just outperforms traditional nuclear weapons on some more local scale. The basic logic behind any bomb is a self-sustaining chain reaction, in which a carrier mjx-container[jax="CHTML"] { line-height: 0; } mjx-con
Existential RiskPhysicsSpeculative Analysis
Research LessWrong Jul 3

On "gendertropes" in dath ilan

By Eliezer Yudkowsky

18 score
AI Analysis

Yudkowsky explains a constructed-language concept from his fictional world about gendered personality and relationship archetypes, framing it as a radial prototype-based category. The essay is primarily a philosophical and linguistic exposition tied to his fiction rather than technical AI research.

I have sometimes been asked with respect to my fiction, "What the hell is a 'gendertrope'?""Gendertrope" is a word from Baseline, the language of the fictional world of dath ilan; it appears in my stories about dath ilan and dath ilani.I have observed that (1) the meaning of "gendertrope" seems hard for many people to grasp initially, and (2) once people do grasp that meaning, they sometimes use the word in ordinary conversation with other people who know it, including those who previously compl
RationalityPhilosophyFiction