Continuing our coverage from Research two days ago on the Anthropic-DoW confrontation, Zvi analyzes the escalating confrontation between the Department of War and Anthropic, where the Pentagon demanded 'unfettered access' to Claude for all lawful military uses or face designation as a supply chain risk or invocation of the Defense Production Act. The piece covers Anthropic's response, broader industry reactions, and the legal and governance implications of government coercion of AI companies.
Category intelligence
Research Briefing — February 28, 2026
16 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's most impactful work spans AI safety research and a landmark governance crisis between frontier labs and the U.S. military.
- Model Incrimination (MATS 9.0 / Neel Nanda) introduces novel methods to diagnose *why* an LLM misbehaves—distinguishing scheming from sycophancy or capability failures—filling a critical gap in safety evaluation
- Unsupervised Elicitation research (MATS/Anthropic) delivers an important negative result: no existing easy-to-hard generalization technique reliably scales oversight to superhuman models across three realistic challenge settings
- The Dawn of AI Scheming provides the most comprehensive survey to date of empirical evidence on deceptive alignment across model organisms
On governance, Zvi's analysis of the Anthropic–Department of War standoff and Sam Altman's memo aligning OpenAI with Anthropic's red lines on autonomous weapons and mass surveillance represent a potential inflection point for military AI policy. New ARENA exercise sets package frontier interpretability topics (attribution graphs, emergent features) into accessible training material. Abram Demski's Coherent Care advances foundational arguments for Updateless Decision Theory, and a community-built RSP version comparison tool aids timely policy scrutiny.
Key Themes
Primary evidence
Top Ranked Signals
Sam Altman says OpenAI shares Anthropic's red lines in Pentagon fight
By Matrice Jacobine
Building on yesterday's News coverage of Anthropic's refusal, Sam Altman circulated an internal memo stating OpenAI will draw the same red lines as Anthropic regarding Pentagon AI use—no mass surveillance or autonomous lethal weapons. This represents a potential industry-wide unified stance that could complicate the Pentagon's efforts to replace Anthropic with another AI provider.
Why Did My Model Do That? Model Incrimination for Diagnosing LLM Misbehavior
By aditya singh
MATS 9.0 research (advised by Neel Nanda) introducing 'model incrimination'—methods to determine whether a model's suspicious behavior stems from scheming, confusion, or mistakes. They build environments where models take concerning actions and use interpretability techniques to investigate the underlying motivations, aiming to help labs distinguish genuine scheming from benign errors.
3 Challenges and 2 Hopes for the Safety of Unsupervised Elicitation
By Callum Canavan
Research from MATS/Anthropic fellowship studying three realistic challenges to unsupervised elicitation and easy-to-hard generalization techniques for steering models on superhuman tasks. They stress-test existing techniques and two new approaches (ensembling, combined methods), finding that no technique reliably overcomes all three challenges.
New ARENA material: 8 exercise sets on alignment science & interpretability
By CallumMcDougall
Callum McDougall announces 8 new ARENA exercise sets covering cutting-edge alignment and interpretability topics including attribution graphs, emergent misalignment, alignment faking, reasoning model interpretability (Thought Anchors), and persona vectors. Each set contains 1-2 days of hands-on material for upskilling researchers.
A comprehensive aggregation of virtually everything currently known about AI scheming (deceptive alignment), covering empirical evidence from model organisms, theoretical arguments, and building toward an informed forecast. Written primarily during autumn 2025, it serves as a reference document for the scheming threat model.
Following News coverage from two days ago of Anthropic's RSP changes, A community member created a side-by-side comparison website for all versions of Anthropic's Responsible Scaling Policy (RSP), making it easy to track how the policy has evolved. The author notes a clear shift from firm commitments to softer reporting language across versions.
Abram Demski presents careful arguments for Updateless Decision Theory (UDT) over naively updateful theories like CDT and EDT, framed through his ongoing tiling theorem research agenda. The essay develops the idea that rational agents should 'care coherently' about outcomes in a way that motivates updateless reasoning, with implications for AI safety foundations.
An essay arguing that safe ASI is achievable by framing alignment as a 'finite game' with identifiable win conditions, written against the backdrop of Anthropic's RSP changes and rapidly advancing model capabilities (Opus 4.6, Codex 5.3). The author attempts to provide an optimistic but grounded framework for why alignment challenges are solvable.
A practitioner's observation that effective 'vibe coding' (AI-assisted programming) converges on skills similar to system design interviews: writing detailed specifications, thinking through edge cases, and making architectural tradeoffs, rather than writing code directly.
A thought experiment exploring how physical systems like a ball rolling downhill can be described using the language of preferences from Optimal Information Structure (OIS) theory. The author argues that composite systems (ball+gravity) exhibit preferences in a technical sense analogous to human preferences.
Announcement for a 7-day intensive AI security bootcamp in Singapore targeting experienced security professionals, covering threat modeling for frontier AI systems, adversarial misuse, loss-of-control risks, and governance interventions. This is the second iteration after a London program in August 2025.