Category intelligence

Research Briefing — February 7, 2026

16 current items analyzed and ranked.

Executive synthesis

Research Summary

The dominant theme is a sharp debate over interpretability-in-the-loop training—using interpretability signals in loss functions. Steven Byrnes offers a rigorous conditional defense of the technique, while a separate post flags Goodfire as actively deploying it, raising safety concerns about what some call 'The Most Forbidden Technique.'

On the practical side, early impressions of Claude Opus 4.6 (released 2026-02-05) highlight its agent swarm mode and notably increased 'drive' in agentic coding tasks. A factorial experiment (n=900, Cohen's d=2.67) demonstrates that prompt imperativeness drastically reduces LLM hedging behavior, with immediate practical implications for prompt engineering.

Key Themes

Mechanistic Interpretability · 4AI Safety & Alignment · 8Agentic AI & Coding Tools · 3AI Benchmarking & Evaluation · 1AI Governance & Policy · 4

Primary evidence

Top Ranked Signals

72 score
AI Analysis

Following yesterday's News coverage of Goodfire AI, Steven Byrnes offers a conditional defense of 'interpretability-in-the-loop training' (using interpretability signals in the loss function), which is widely considered dangerous because it could train models to obfuscate their reasoning. He argues there may be narrow conditions where the approach is valid, pushing back against the blanket prohibition.

Let’s call “interpretability-in-the-loop training” the idea of running a learning algorithm that involves an inscrutable trained model, and there’s some kind of interpretability system feeding into the loss function / reward function.Interpretability-in-the-loop training has a very bad rap (and rightly so). Here’s Yudkowsky 2022:When you explicitly optimize against a detector of unaligned thoughts, you're partially optimizing for more aligned thoughts, and partially optimizing for unaligned
AI SafetyMechanistic InterpretabilityAlignmentTraining Methodology
68 score
AI Analysis

Introduces 'Meta-Autointerp,' a method using pretrained SAEs alongside LLM-summarizer methods to interpret multi-agent RL training runs in the game Diplomacy. The approach discovers fine-grained behavioral patterns and, when discovered features are added to an untrained agent's system prompt, improves performance by 14.2%.

TLDR; SAEs can complement and enhance LLM as a Judge scalable oversight for uncovering hypotheses over large datasets of LLM outputspaperAbstractLarge language models (LLMs) are increasingly trained in long-horizon, multi-agent environments, making it difficult to understand how behavior changes over training. We apply pretrained SAEs, alongside LLM-summarizer methods, to analyze reinforcement learning training runs from Full-Press Diplomacy, a long-horizon multi-player strategy game. We introdu
Mechanistic InterpretabilityMulti-Agent RLScalable OversightSAE Research
Research LessWrong Feb 6

AI benchmarking has a Y-axis problem

By Lizka

58 score
AI Analysis

Argues that AI benchmark scores lack natural units, making it misleading to plot them over time and draw conclusions about acceleration, inflection points, or trends. Benchmark scores are 'funhouse-mirror projections' of true capability that compress and stretch different capability regions arbitrarily.

TLDR: People plot benchmark scores over time and then do math on them, looking for speed-ups & inflection points, interpreting slopes, or extending apparent trends. But that math doesn’t actually tell you anything real unless the scores have natural units. Most don’t.Think of benchmark scores as funhouse-mirror projections of “true” capability-space, which stretch some regions and compress others by assigning warped scores for how much accomplishing that task counts in units of “AI progress”
AI BenchmarkingMethodologyAI Progress Measurement
Research LessWrong Feb 6

Robust Finite Policies are Nontrivially Structured

By Winter Cross

55 score
AI Analysis

A formal result showing that policies modeled as deterministic finite automata must share nontrivial structural features if they meet certain robustness criteria. This is a step toward the 'agent structure problem'—the conjecture that agent-like behavior implies agent-like internal structure.

This post was created during the Dovetail Research Fellowship. Thanks to Alex, Alfred,  everyone who read and commented on the draft, and everyone else in the fellowship for their ideas and discussions.OverviewThe proof detailed in this post was motivated by a desire to take a step towards solving the agent structure problem, which is the conjecture that a system which exhibits agent-like behavior must have agent-like structure. Our goal was to describe a scenario where something concrete a
Agent FoundationsAI SafetyFormal MethodsAlignment Theory
Research LessWrong Feb 5

Goodfire and Training on Interpretability

By Satya Benson

55 score
AI Analysis

Following yesterday's News coverage of Goodfire AI, A brief post flagging that Goodfire is actively pursuing 'training on interpretability'—using interpretability signals in the training loop—which the AI safety community has repeatedly warned against as 'The Most Forbidden Technique.' Asks for community evaluation of Goodfire's claimed risk management.

Goodfire wrote Intentionally designing the future of AI about training on interpretability.This seems like an instance of The Most Forbidden Technique which has been warned against over and over - optimization pressure on interpretability technique [T] eventually degrades [T].Goodfire claims they are aware of the associated risks and managing those risks.Are they properly managing those risks? I would love to get your thoughts on this.
AI SafetyMechanistic InterpretabilityTraining MethodologyAlignment
Research LessWrong Feb 6

Claude Code #4: From The Before Times

By Zvi

52 score
AI Analysis

Following yesterday's News coverage of the Opus 4.6 release, Zvi's fourth installment covering Claude Code best practices, written just before the Claude Opus 4.6 and agent swarm announcements. Covers practical techniques for agentic coding workflows, task decomposition, verification strategies, and tips for getting maximum value from AI coding assistants.

Claude Opus 4.6 and agent swarms were announced yesterday. That’s some big upgrades for Claude Code. OpenAI, the competition, offered us GPT-5.3-Codex, and this week gave us an app form of Codex that already has a million active users. That’s all very exciting, and next week is going to be about covering that. This post is about all the cool things that happened before that, which we will be building upon now that capabilities have further advanced. This if from Before Times. Almost all of it st
AI Coding ToolsClaude CodeHuman-AI CollaborationAgentic AI
Research LessWrong Feb 5

Claude Opus 4.6 is Driven

By HunterJay

48 score
AI Analysis

Following yesterday's News coverage of the Opus 4.6 release, First-day impressions of Claude Opus 4.6 and its new agent swarm/teams mode in Claude Code. The author tested it on a production codebase and reports the model is notably 'driven' and goal-oriented, with the agent swarm mode showing promise for parallel feature development.

Claude is driven to achieve it's goals, possessed by a demon, and raring to jump into danger. These are my impressions from the first day of usage. Epistemic status: personal observations and quotes from more reliable sources.____Today Claude Opus 4.6 was launched along with an update to Claude Code which enabled a ‘teams’ mode (also known as an Agent Swarm). The mode sets up multiple agents to run in parallel with a supervisor, and are provided with methods of communicating between themsel
Claude Opus 4.6Agentic AIAI Coding ToolsModel Evaluation
45 score
AI Analysis

A factorial experiment (n=900) showing that prompt imperativeness drastically reduces hedging in LLMs, with a very large effect size (Cohen's d=2.67). Tested across three models, two question types, and three imperativeness levels; Claude hedged the most.

ForewordIf you've ever noticed LLMs hedging considerably more when you ask them subjective questions, it's not a fluke. I ran a 3x2x3 factorial experiment (n=900) to quantify how much prompt phrasing (alongside question type and model type) shifts hedging across differing imperativeness levels. The effect sizes were larger than I expected.To nobody's surprise, Claude hedged the most (by a fairly wide margin). It also decided to meta-analyze its own response then critiqued its own compliance in a
LLM BehaviorPrompt EngineeringEmpirical AI Research
42 score
AI Analysis

Critiques Dario Amodei's 'The Adolescence of Technology' essay as delegitimizing AI x-risk concerns. Argues that Anthropic's framing as the 'responsible racer' actually normalizes the dangerous race toward superintelligence more effectively than OpenAI or xAI do.

My beef with AnthropicI've long felt that while Anthropic is the most safety-conscious of the frontier AI companies, they're also the most hypocritical enablers of the whole reckless enterprise. By framing themselves as the "good sport" in the race, the one who's encouraging everyone else to "race them to the top", the one who's making sacrifices on the margin so as to be the "best of the worst" — they're actually the ones broadcasting the most powerful signal that racing toward the superintelli
AI SafetyAI GovernanceAnthropicAI Policy
Research LessWrong Feb 6

Spectral Signatures of Gradual Disempowerment

By Jonas Hallgren

38 score
AI Analysis

Proposes using spectral graph theory metrics (spectral gap, Fiedler vector, eigenvalue distributions) as cross-domain measures for tracking gradual human disempowerment as AI systems enter coordination systems like markets, networks, and governance institutions.

TL;DRAI disempowerment operates across markets, networks, and governance simultaneously, but our analytical tools don't cross those boundaries. We propose spectral graph metrics—spectral gap, Fiedler vector, eigenvalue distribution—as computable, cross-domain measures for tracking how the balance of influence shifts when AI enters coordination systems, and identify three specific quantities to monitor for AI governance.IntroductionAI systems are changing how society coordinates — across markets,
AI GovernanceAI SafetyHuman DisempowermentNetwork Analysis
Research LessWrong Feb 6

Strategy of von Neumann and strategy of Rosenbergs

By avturchin

25 score
AI Analysis

An essay drawing an analogy between Cold War nuclear proliferation strategies (von Neumann's 'strike first' vs. the Rosenbergs' 'share the technology') and potential approaches to AI development concentration. The author explores whether distributing powerful AI capabilities widely could be a stability strategy analogous to nuclear proliferation.

This is not a call for espionage, but an analysis of another strategyVon Neumann's strategy for solving the problem of global nuclear weapons proliferation is widely known - strike tomorrow. That is, conquer the entire world by exploiting that brief window when only one side possesses nuclear weapons. This idea is popular among American readers, partly because personal interests for the US correlate with this strategy: It would be good for the world and for us. (I will not discuss here whether v
AI GovernanceAI StrategyExistential Risk
Research LessWrong Feb 5

Why ASI Might Preserve Its Progenitors

By Luke J. Dawes

22 score
AI Analysis

Argues that even a misaligned ASI might have instrumental reasons to preserve humanity, particularly if it assigns non-negligible probability to future observers (aliens, simulators) who might judge its behavior. Draws on decision theory and Hanson's grabby aliens model.

SummaryEven a misaligned Earth-originating artificial superintelligence (ASI) could have instrumentally rational reasons, under multiple decision theories, to preserve rather than destroy humanity. This would depend on the ASI assigning non-negligible probability to the existence of future observers (e.g. intelligent aliens, their ASIs, or simulators).IntroductionWe may build AI so powerful that it is better than us at everything: call this artificial superintelligence (ASI). We may build one AS
AI SafetyExistential RiskDecision TheoryASI Speculation