Category intelligence

Research Briefing — February 28, 2026

16 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's most impactful work spans AI safety research and a landmark governance crisis between frontier labs and the U.S. military.

  • Model Incrimination (MATS 9.0 / Neel Nanda) introduces novel methods to diagnose *why* an LLM misbehaves—distinguishing scheming from sycophancy or capability failures—filling a critical gap in safety evaluation
  • Unsupervised Elicitation research (MATS/Anthropic) delivers an important negative result: no existing easy-to-hard generalization technique reliably scales oversight to superhuman models across three realistic challenge settings
  • The Dawn of AI Scheming provides the most comprehensive survey to date of empirical evidence on deceptive alignment across model organisms

On governance, Zvi's analysis of the Anthropic–Department of War standoff and Sam Altman's memo aligning OpenAI with Anthropic's red lines on autonomous weapons and mass surveillance represent a potential inflection point for military AI policy. New ARENA exercise sets package frontier interpretability topics (attribution graphs, emergent features) into accessible training material. Abram Demski's Coherent Care advances foundational arguments for Updateless Decision Theory, and a community-built RSP version comparison tool aids timely policy scrutiny.

Key Themes

AI Governance & Pentagon Standoff · 4AI Safety & Alignment Research · 6Interpretability & Mechanistic Understanding · 3AI Education & Community Building · 3Agent Foundations & Decision Theory · 2

Primary evidence

Top Ranked Signals

Research LessWrong Feb 27

Anthropic and the DoW: Anthropic Responds

By Zvi

78 score
AI Analysis

Continuing our coverage from Research two days ago on the Anthropic-DoW confrontation, Zvi analyzes the escalating confrontation between the Department of War and Anthropic, where the Pentagon demanded 'unfettered access' to Claude for all lawful military uses or face designation as a supply chain risk or invocation of the Defense Production Act. The piece covers Anthropic's response, broader industry reactions, and the legal and governance implications of government coercion of AI companies.

The Department of War gave Anthropic until 5:01pm on Friday the 27th to either give the Pentagon ‘unfettered access’ to Claude for ‘all lawful uses,’ or else. With the ‘or else’ being not the sensible ‘okay we will cancel the contract then’ but also expanding to either being designated a supply chain risk or having the government invoke the Defense Production Act. It is perfectly legitimate for the Department of War to decide that it does not wish to continue on Anthropic’s terms, and that it wi
AI GovernanceAI PolicyNational SecurityAI Safety
Research LessWrong Feb 27

Sam Altman says OpenAI shares Anthropic's red lines in Pentagon fight

By Matrice Jacobine

75 score
AI Analysis

Building on yesterday's News coverage of Anthropic's refusal, Sam Altman circulated an internal memo stating OpenAI will draw the same red lines as Anthropic regarding Pentagon AI use—no mass surveillance or autonomous lethal weapons. This represents a potential industry-wide unified stance that could complicate the Pentagon's efforts to replace Anthropic with another AI provider.

OpenAI CEO Sam Altman wrote in a memo to staff that he will draw the same red lines that sparked a high-stakes fight between rival Anthropic and the Pentagon: no AI for mass surveillance or autonomous lethal weapons.Why it matters: If other leading firms like Google follow suit, this could massively complicate the Pentagon's efforts to replace Anthropic's Claude, which was the first model integrated into the military's most sensitive work.It would also be the first time the nation's top AI leade
AI GovernanceAI PolicyNational SecurityAI Ethics
72 score
AI Analysis

MATS 9.0 research (advised by Neel Nanda) introducing 'model incrimination'—methods to determine whether a model's suspicious behavior stems from scheming, confusion, or mistakes. They build environments where models take concerning actions and use interpretability techniques to investigate the underlying motivations, aiming to help labs distinguish genuine scheming from benign errors.

Authors: Aditya Singh*, Gerson Kroiz*, Senthooran Rajamanoharan, Neel NandaAditya and Gerson are co-first authors. This work was conducted during MATS 9.0 and was advised by Senthooran Rajamanoharan and Neel Nanda.MotivationImagine that a frontier lab’s coding agent has been caught putting a bug in the key code for monitoring what that agent does. Naively, this seems like a clear smoking gun that the agent is scheming. But LLMs often do weird things; they could easily just be confused, or have m
AI SafetyInterpretabilityAI SchemingAlignment
Research LessWrong Feb 27

3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation

By Callum Canavan

68 score
AI Analysis

Research from MATS/Anthropic fellowship studying three realistic challenges to unsupervised elicitation and easy-to-hard generalization techniques for steering models on superhuman tasks. They stress-test existing techniques and two new approaches (ensembling, combined methods), finding that no technique reliably overcomes all three challenges.

Authors: Callum Canavan*, Aditya Shrivastava*, Allison Qi, Jonathan Michala, Fabien Roger(*Equal contributions, alphabetical)tl;dr: We study 3 realistic challenges to the safety of unsupervised elicitation and easy-to-hard generalization techniques, which aim to steer models on tasks which are beyond human supervision. We create datasets to test the robustness of methods against these challenges. We stress-test existing techniques on them along with new methods relying on 2 hopes: ensembling and
AI SafetyScalable OversightAlignmentElicitation
62 score
AI Analysis

Callum McDougall announces 8 new ARENA exercise sets covering cutting-edge alignment and interpretability topics including attribution graphs, emergent misalignment, alignment faking, reasoning model interpretability (Thought Anchors), and persona vectors. Each set contains 1-2 days of hands-on material for upskilling researchers.

TLDRThis is a post announcing a lot of new ARENA material I've been working on for a while, which is now available for study here (currently on the alignment-science branch, but planned to be merged into main this Sunday).There's a set of exercises (each one contains about 1-2 days of material) on the following topics:Linear Probes (replication of the "Geometry of Truth" paper, plus Apollo's "Probing for Deception" work)Activation Oracles (based around this demo notebook, with additional exercis
AI SafetyInterpretabilityAI EducationAlignment
Research LessWrong Feb 27

The Dawn of AI Scheming

By Alvin Ånestrand

55 score
AI Analysis

A comprehensive aggregation of virtually everything currently known about AI scheming (deceptive alignment), covering empirical evidence from model organisms, theoretical arguments, and building toward an informed forecast. Written primarily during autumn 2025, it serves as a reference document for the scheming threat model.

This article aggregates virtually everything currently known about AI scheming, then builds toward an informed forecast.How to read this article: After reading the introduction to understand the article’s scope and structure, I recommend moving directly to the Overview and Forecast sections, and read the background sections as needed for context. The article is long, so feel free to prioritize the background material you find most relevant.When you see the terms “AI” or “model”, they usually ref
AI SafetyAI SchemingDeceptive AlignmentAI Alignment
Research LessWrong Feb 27

Side by Side Comparison of RSP Versions

By Corm

45 score
AI Analysis

Following News coverage from two days ago of Anthropic's RSP changes, A community member created a side-by-side comparison website for all versions of Anthropic's Responsible Scaling Policy (RSP), making it easy to track how the policy has evolved. The author notes a clear shift from firm commitments to softer reporting language across versions.

With all of the discussion about changes to Anthropic's Responsible Scaling Policy, I figured actually reading through all of them in one go would be helpful. I wanted to easily compare sections side by side, so I made a quick website which you can find here.It took me a little over an hour to read through all of them and it seemed net worth it. It's really clear how the tone has shifted, going from commitments (of which many have definitely been broken) to something more along the lines of repo
AI SafetyAI GovernanceResponsible Scaling
Research LessWrong Feb 27

Coherent Care

By abramdemski

40 score
AI Analysis

Abram Demski presents careful arguments for Updateless Decision Theory (UDT) over naively updateful theories like CDT and EDT, framed through his ongoing tiling theorem research agenda. The essay develops the idea that rational agents should 'care coherently' about outcomes in a way that motivates updateless reasoning, with implications for AI safety foundations.

I've been trying to gather my thoughts for my next tiling theorem (agenda write-up here; first paper; second paper; recent project update). I have a lot of ideas for how to improve upon my work so far, and trying to narrow them down to an achievable next step has been difficult. However, my mind keeps returning to specific friends who are not yet convinced of Updateless Decision Theory (UDT).I am not out to argue that UDT is the perfect decision theory; see eg here and here. However, I strongly
Agent FoundationsDecision TheoryAI Safety
Research LessWrong Feb 27

Safe ASI Is Achievable: The Finite Game Argument

By Lester Leong

40 score
AI Analysis

An essay arguing that safe ASI is achievable by framing alignment as a 'finite game' with identifiable win conditions, written against the backdrop of Anthropic's RSP changes and rapidly advancing model capabilities (Opus 4.6, Codex 5.3). The author attempts to provide an optimistic but grounded framework for why alignment challenges are solvable.

A few days ago, Anthropic dropped the central pledge of its Responsible Scaling Policy, the promise that it would never train an AI system unless it could guarantee in advance that its safety measures were adequate. The stated reason: unilateral safety commitments don't make sense when competitors are racing ahead without them. METR's Chris Painter, who reviewed an early draft, put it bluntly: Anthropic "believes it needs to shift into triage mode with its safety plans, because methods to assess
AI SafetyAI AlignmentExistential Risk
Research LessWrong Feb 26

Vibe Coding is a System Design Interview

By Brendan Long

25 score
AI Analysis

A practitioner's observation that effective 'vibe coding' (AI-assisted programming) converges on skills similar to system design interviews: writing detailed specifications, thinking through edge cases, and making architectural tradeoffs, rather than writing code directly.

I've been working on two fairly large vibe-coded apps, and my process has converged on:Write a GitHub issue(If complicated enough) tell an agent to make a plan and then update the issueHave another agent read the issue and implement itAs the features get more complicated, I spend more and more time on step (1), and I'm finding that just taking the time to write a detailed enough issue is 90% of the work (and if I have a problem, going back and writing a much more detailed issue usually fixes it)
AI-Assisted DevelopmentSoftware Engineering
Research LessWrong Feb 27

Ball+Gravity has a "Downhill" Preference

By TristanTrim

15 score
AI Analysis

A thought experiment exploring how physical systems like a ball rolling downhill can be described using the language of preferences from Optimal Information Structure (OIS) theory. The author argues that composite systems (ball+gravity) exhibit preferences in a technical sense analogous to human preferences.

[epistemic status: This is a rambling thought experiment with the goal of clarifying my ontological understanding of "agent foundations" type stuff. Scroll to the bottom for the resulting two "interesting focuses of confusion".]Returning to the "ball rolls down a hill" example of Alex Altair's My research agenda in agent foundations, I think there's three important objects here: The ball, the hill, and gravity. Friction and inertia and other physics are also doing important things here in a real
Agent FoundationsAI Safety
Research LessWrong Feb 27

AI Security Bootcamp Singapore - Call for Applications

By Pranav Gade

15 score
AI Analysis

Announcement for a 7-day intensive AI security bootcamp in Singapore targeting experienced security professionals, covering threat modeling for frontier AI systems, adversarial misuse, loss-of-control risks, and governance interventions. This is the second iteration after a London program in August 2025.

tl;drWe're running a 7-day intensive AI security program in Singapore for experienced security professionals who want to upskill on securing frontier AI systems. This is the second iteration of AISB - the first ran in London in August 2025. Accommodation and programme costs are covered, and limited travel support is available if you need financial assistance.Apply by March 15, 2026.Why this program existsAs AI systems become more capable and integrated into critical infrastructure, new attack su
AI SecurityAI SafetyAI Education