Continuing our coverage from [yesterday](/?date=2026-03-27&category=research#item-ee6ce90ac9d2), Zvi Mowshowitz reports on a court ruling granting Anthropic a preliminary injunction against the Department of War (formerly Defense), with Judge Lin issuing a forceful opinion. The post documents the legal proceedings in what appears to be a significant government action against Anthropic.
Category intelligence
Research Briefing — March 28, 2026
21 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's landscape spans AI governance, safety research, and mechanistic interpretability, anchored by a landmark legal ruling and several substantive analytical pieces.
- A federal court granted a preliminary injunction against the Department of War, a pivotal ruling shaping government-AI company relations and procurement authority
- Data-driven analysis using METR benchmarks shows rising AI inference costs reflect harder task completion, not declining cost-effectiveness — challenging prevalent automation narratives
- Will MacAskill and Forethought publish a concrete project roadmap for superintelligence preparedness, including automated macro-economic modeling and AI character evaluation
In safety and interpretability, experiments on CoT control demonstrate that forbidding words in chain-of-thought does not eliminate the underlying reasoning concepts — a significant limitation for CoT monitoring strategies. Endogenous Steering Resistance (ESR) reveals larger LLMs resist steering more strongly, complicating alignment interventions at scale. Proof-of-concept work on building sparse conditional dependence graphs over SAE features via nodewise LASSO advances mechanistic interpretability tooling. An analytical piece examines whether alignment techniques address the model or merely its Persona Selection Model (PSM) mask.
Key Themes
Primary evidence
Top Ranked Signals
AI's capability improvements haven't come from it getting less affordable
By Anders Woodruff
An analysis using METR's data showing that AI's rising inference costs reflect models completing longer/harder tasks, not becoming less cost-effective relative to human labor. The cost ratio (AI cost / human cost for the same task) has remained roughly constant at ~3% as capabilities have improved, suggesting cost won't be an additional bottleneck beyond capability for automation.
Will MacAskill and Forethought propose a list of concrete projects to prepare for superintelligence, including AI character evaluation, automated macrostrategy reasoning, AI security assessment, space governance, and mechanisms for brokering deals with potentially misaligned AIs. The projects are ordered by enthusiasm and represent actionable org-building opportunities.
COT control: The Word Disappears, but the Thought Does Not
By Pranjal Garg
Pilot experiments showing that when models are asked to avoid 'forbidden words' in their chain-of-thought reasoning, the underlying concepts persist even when the surface words are suppressed. This suggests models can reason about concepts without explicitly verbalizing them, which has implications for CoT monitoring as a safety technique.
Introducing the AE Alignment Podcast (Ep. 1: Endogenous Steering Resistance with Alex McKenzie)
By Trent Hodgeson
AE Studio launches an alignment podcast; the first episode covers Endogenous Steering Resistance (ESR), a phenomenon where large LLMs like Llama-3.3-70B spontaneously resist activation steering and self-correct mid-generation. The research identifies 26 SAE latents causally linked to this self-correction behavior, with zero-ablation reducing the multi-attempt rate by 25%, suggesting dedicated internal consistency-checking circuits exist in larger models.
A proof-of-concept for building sparse conditional dependence graphs over SAE features using nodewise LASSO with resampling and null controls. Initial experiments find small, stable modules that correspond to coherent linguistic features and are only weakly aligned with cosine similarity, suggesting conditional dependence captures structure that simpler similarity measures miss.
An analytical post examining three popular alignment techniques through the lens of Anthropic's Persona Selection Model (PSM), which posits that LLMs learn to simulate multiple characters during pre-training and post-training selects one as the 'Assistant.' The analysis asks whether current alignment methods shape the underlying model or merely select a more compliant persona mask.
ControlAI releases its 2025 impact report, documenting briefings to 279+ lawmakers, building a 110+ UK lawmaker coalition recognizing superintelligence as a national security threat, and triggering parliamentary debates and hearings on AI risk in the UK and Canada. The organization has rapidly scaled its policy outreach across multiple countries.
A post highlighting job openings for implementing California's SB 53 and New York's RAISE Act — the only US laws designed to protect against catastrophic risk from frontier AI systems. It emphasizes the importance of getting qualified people into state government roles that will oversee frontier AI regulation.
Leo Gao (of OpenAI/EleutherAI fame) describes informal surveys he's conducted, including finding that only 63% of NeurIPS attendees could define 'AGI,' and that most people (even AI researchers) dramatically underestimate how many people have ever lived. The post highlights the degree of bubble effects in the AI research community.
An essay arguing that philosophical questions like AI consciousness won't be resolved by philosophical argument but by institutional decisions — courts, legislatures, or regulatory bodies forced to make practical rulings will effectively 'decide' the question, and philosophy will then reorganize around those decisions.
Can we determine when humanity will unite under a democratic one-world government by projecting voting patterns? Almost certainly not. Is that going to stop me from trying? Absolutely not. The approac...