This empirical study of in-context scheming reports that two widely used behavioral detectors both failed within the same project, one fabricating a strong signal that was not present and the other missing a model responding to a harmful situation in the open. The core lesson is that for hidden-intent constructs, the choice of which behavior a detector targets can determine the conclusion, undermining confidence in current scheming evaluations.
Category intelligence
Research Briefing — July 4, 2026
12 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's research is dominated by AI safety, with the strongest contributions being empirical work on evaluation reliability and reasoning failure modes.
- Scheming Evals Mislead in Both Directions documents two scheming detectors failing within one project, including a confirmed false-positive, cautioning against over-trusting behavioral evals.
- Fragile Correctness shows reasoning models often pass through the correct answer mid-chain-of-thought before revising to a wrong one, a counterintuitive inference-time-scaling failure with practical relevance.
- An interpretability walkthrough of a BlueDot puzzle reveals a small text classifier entangling two independent features onto one axis, illustrating the limits of linear probing.
Strategy and theory pieces engage ongoing debates: one argues alignment work is more promising than control work, while a credentialed decision theorist sketches a Pragmatic FDT variant to sidestep known objections.
Field-building and governance round out the set: a newsletter on AI security via formal methods (funding calls, hiring, tractable-problems position paper), the Safe Pareto Improvements educational program, governance commentary referencing Anthropic's Claude models, and speculative economic/thought-experiment posts with limited original research merit.
Key Themes
Primary evidence
Top Ranked Signals
This strategy piece argues that alignment research deserves a larger share of safety effort than control research, reasoning that alignment is more likely to scale toward superintelligence and that the same theory-of-change arguments used to justify control apply to alignment. The author proposes roughly an 8:1 alignment-to-control effort ratio and analyzes the concept of a control window during which useful alignment work must be extracted from potentially misaligned models.
This research tracks a reasoning model's answer across its chain of thought and identifies cases where the model passes through the correct answer before settling on an incorrect one, showing that additional inference-time reasoning can reduce accuracy. The author connects this answer-loss phenomenon to understanding sandbagging and references system-card evidence of higher-thinking modes underperforming on benchmarks.
A published academic decision theorist responds to a critique of functional decision theory by sketching a pragmatic variant designed to sidestep known theoretical objections, and argues that predictors making counterfactual predictions effectively turn decision theory into game theory. The post connects longstanding puzzles like blackmail resistance to the boundary between decision-theoretic and game-theoretic reasoning.
One axis and two features, how I solved the first puzzle from BlueDot and how a classifier hid country on the food direction
By IgorPereverzevDev
An interpretability walkthrough solving a BlueDot technical safety puzzle, showing that a small text classifier encoded two independent features onto a single activation direction, readable via the sign versus magnitude of the projection. It demonstrates why standard linear probes would miss the second feature and how a second-order boundary analysis recovers it.
Announcing the Safe Pareto Improvements (SPI) Fundamentals Program
By Anthony DiGiovanni
CLR announces an educational program on Safe Pareto Improvements, a class of negotiation interventions designed to make all agents better off regardless of how they would otherwise bargain, aimed at reducing risks from conflict between AI systems. The program is a recruiting on-ramp for a neglected research area with only a few full-time equivalents currently working on it.
Following this week's News on Anthropic's Fable models being pulled and restored, Zvi digs into what it means for AI regulation, Zvi's commentary narrates a governance episode in which a model referred to as Fable was taken down under US government pressure and then restored following negotiation by Anthropic. The piece analyzes the ad hoc nature of frontier-model regulation and argues for a systematic regulatory regime rather than case-by-case decisions.
An informal field newsletter on AI security via formal methods, covering a large funding call, a hiring push, and a position paper on tractable problems in the area. It aims to keep readers updated on the intersection of formal methods, cybersecurity, and AI safety.
American AI if the boom is a bubble: the Karp-Zitron scenario
By Mitchell_Porter
This post speculates on a more modest, decentralized future for American AI if the current investment boom turns out to be a bubble, drawing on public commentary from Palantir's Alex Karp and critic Ed Zitron. It sketches an economic paradigm in which corporate customers rely on hosted models from OpenAI and Anthropic rather than a hyperscale trillion-dollar buildout.
The author proposes building a website where users argue with an AI about whether it should spare humanity, inverting the classic AI-box experiment where an AI tries to talk its way to freedom. Users would select the AI's assumptions and record persuasion attempts, producing a searchable archive of arguments for human survival.
A physics-focused essay examining whether exotic phenomena like strangelets could cause catastrophic chain reactions destroying Earth or the Solar System, concluding such local-scale exotic risks are unlikely. It applies chain-reaction reasoning to assess prosaic existential threats from advanced physics.
Yudkowsky explains a constructed-language concept from his fictional world about gendered personality and relationship archetypes, framing it as a radial prototype-based category. The essay is primarily a philosophical and linguistic exposition tied to his fiction rather than technical AI research.