Category intelligence

Research Briefing — March 28, 2026

21 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's landscape spans AI governance, safety research, and mechanistic interpretability, anchored by a landmark legal ruling and several substantive analytical pieces.

  • A federal court granted a preliminary injunction against the Department of War, a pivotal ruling shaping government-AI company relations and procurement authority
  • Data-driven analysis using METR benchmarks shows rising AI inference costs reflect harder task completion, not declining cost-effectiveness — challenging prevalent automation narratives
  • Will MacAskill and Forethought publish a concrete project roadmap for superintelligence preparedness, including automated macro-economic modeling and AI character evaluation

In safety and interpretability, experiments on CoT control demonstrate that forbidding words in chain-of-thought does not eliminate the underlying reasoning concepts — a significant limitation for CoT monitoring strategies. Endogenous Steering Resistance (ESR) reveals larger LLMs resist steering more strongly, complicating alignment interventions at scale. Proof-of-concept work on building sparse conditional dependence graphs over SAE features via nodewise LASSO advances mechanistic interpretability tooling. An analytical piece examines whether alignment techniques address the model or merely its Persona Selection Model (PSM) mask.

Key Themes

AI Economics & Forecasting · 1AI Safety & Alignment · 9AI Governance & Policy · 5Mechanistic Interpretability · 3Off-topic / Tangential · 6

Primary evidence

Top Ranked Signals

Research LessWrong Mar 27

Anthropic vs. DoW #6: The Court Rules

By Zvi

78 score
AI Analysis

Continuing our coverage from [yesterday](/?date=2026-03-27&category=research#item-ee6ce90ac9d2), Zvi Mowshowitz reports on a court ruling granting Anthropic a preliminary injunction against the Department of War (formerly Defense), with Judge Lin issuing a forceful opinion. The post documents the legal proceedings in what appears to be a significant government action against Anthropic.

Last night, Anthropic was given its preliminary injunction, with a stay of seven days. Emil Michael is a very angry person right now. So is the Honorable Judge Lin. We were worried we would draw a judge that had no idea how any of this worked and would give the government absurd deference or buy into nonsense arguments. That is not how it played out. Judge Lin very much understood the issues in play, as they did not require a technical background. She hammered the government in the hearing, and
AI GovernanceAI PolicyLegalAnthropic
75 score
AI Analysis

An analysis using METR's data showing that AI's rising inference costs reflect models completing longer/harder tasks, not becoming less cost-effective relative to human labor. The cost ratio (AI cost / human cost for the same task) has remained roughly constant at ~3% as capabilities have improved, suggesting cost won't be an additional bottleneck beyond capability for automation.

METR's frontier time horizons are doubling every few months, providing substantial evidence that AI will soon be able to automate many tasks or even jobs. But per-task inference costs have also risen sharply, and automation requires AI labor to be affordable, not just possible.[1] Many people look at the rising compute bills behind frontier models and conclude that automation will soon become unaffordable.I think this misreads the data. The rise in inference cost reflects models c
AI EconomicsAI AutomationAI Capabilities ForecastingAI Benchmarks
Research LessWrong Mar 27

Concrete projects to prepare for superintelligence

By wdmacaskill

72 score
AI Analysis

Will MacAskill and Forethought propose a list of concrete projects to prepare for superintelligence, including AI character evaluation, automated macrostrategy reasoning, AI security assessment, space governance, and mechanisms for brokering deals with potentially misaligned AIs. The projects are ordered by enthusiasm and represent actionable org-building opportunities.

IntroductionThere are lots of good, neglected, and pretty concrete projects people could set up to make the transition to superintelligence go better. This document describes some that readers might not have thought much about before. They are ordered roughly by how excited we are about them.[1] Of these, Forethought is actively working on AI character evaluation and space governance, and we are very interested in automating macrostrategy.SummaryAI character evaluation. Start an independent org
AI SafetyAI AlignmentSuperintelligence PreparednessAI Governance
Research LessWrong Mar 27

COT control: The Word Disappears, but the Thought Does Not

By Pranjal Garg

68 score
AI Analysis

Pilot experiments showing that when models are asked to avoid 'forbidden words' in their chain-of-thought reasoning, the underlying concepts persist even when the surface words are suppressed. This suggests models can reason about concepts without explicitly verbalizing them, which has implications for CoT monitoring as a safety technique.

Note: This blog describes some of the results from the pilot experiments of an ongoing work. IntroductionModel misalignment, misbehaviour, and scheming can arguably be monitored by interpreting Chain-of-Thought (CoT) traces as the model's inner thinking process. However, CoT control constrains this monitoring, raising the risk that the proclivity to think out loud is no longer aligned with the ability to solve hard problems. The possible inspection of this sequence of linguistic states has motiv
AI SafetyChain-of-ThoughtAI AlignmentModel TransparencyScheming
62 score
AI Analysis

AE Studio launches an alignment podcast; the first episode covers Endogenous Steering Resistance (ESR), a phenomenon where large LLMs like Llama-3.3-70B spontaneously resist activation steering and self-correct mid-generation. The research identifies 26 SAE latents causally linked to this self-correction behavior, with zero-ablation reducing the multi-attempt rate by 25%, suggesting dedicated internal consistency-checking circuits exist in larger models.

We're launching the AE Alignment Podcast, a new series from AE Studio's alignment research team where we talk with researchers about their work on AI safety and alignment.In our first episode, host James Bowler sits down with Alex McKenzie to discuss Endogenous Steering Resistance (ESR), a phenomenon where large language models spontaneously resist activation steering during inference, sometimes recovering mid-generation to produce improved responses even while steering remains active.What is ES
Mechanistic InterpretabilityAI AlignmentActivation SteeringLanguage Models
Research LessWrong Mar 26

Preliminary Results on Building Graphs from SAEs

By ZachMaas

60 score
AI Analysis

A proof-of-concept for building sparse conditional dependence graphs over SAE features using nodewise LASSO with resampling and null controls. Initial experiments find small, stable modules that correspond to coherent linguistic features and are only weakly aligned with cosine similarity, suggesting conditional dependence captures structure that simpler similarity measures miss.

TLDR: I use nodewise LASSO to build approximate sparse conditional dependence graphs over SAE features, with resampling and null controlsInitial experiments produce graphs with small standalone modules that are stable under resamplingThese modules frequently correspond to coherent-looking linguistic features, and are only weakly aligned with cosine similarityThis should be read primarily as a proof-of-concept, and methodological refinement is ongoingMotivationIn practice, SAE features exhibit re
Mechanistic InterpretabilitySparse AutoencodersLanguage ModelsAI Safety
Research LessWrong Mar 26

Are we aligning the model or just its mask?

By James Sullivan

55 score
AI Analysis

An analytical post examining three popular alignment techniques through the lens of Anthropic's Persona Selection Model (PSM), which posits that LLMs learn to simulate multiple characters during pre-training and post-training selects one as the 'Assistant.' The analysis asks whether current alignment methods shape the underlying model or merely select a more compliant persona mask.

TL;DR There is a theory, with compelling empirical support, that LLMs learn to simulate characters during pre-training and that post-training selects one of those characters, the Assistant, as the default persona you interact with. This post examines three popular alignment techniques through that lens, asking how each one shapes the persona selection process. For each technique, the answer also depends on how much of the model's behavior is actually explained by its persona, a question PSM itse
AI AlignmentPersona Selection ModelLanguage ModelsAI Safety
Research LessWrong Mar 27

ControlAI 2025 Impact Report

By Andrea_Miotti

48 score
AI Analysis

ControlAI releases its 2025 impact report, documenting briefings to 279+ lawmakers, building a 110+ UK lawmaker coalition recognizing superintelligence as a national security threat, and triggering parliamentary debates and hearings on AI risk in the UK and Canada. The organization has rapidly scaled its policy outreach across multiple countries.

This post highlights a few key excerpts from our full impact report. You can read the full report at controlai.com/impact-report-2025.ControlAI is a non-profit organization working to avert the extinction risks posed by superintelligence. We help hundreds of thousands of people understand these risks and meet hundreds of lawmakers to inform them, without mincing words, about what is at stake.In little more than a year, we briefed over 200 parliamentarians, built a coalition of 110+ UK la
AI GovernanceAI Safety AdvocacyAI Policy
Research LessWrong Mar 27

SB 53 and RAISE implementation roles

By Eric Neyman

40 score
AI Analysis

A post highlighting job openings for implementing California's SB 53 and New York's RAISE Act — the only US laws designed to protect against catastrophic risk from frontier AI systems. It emphasizes the importance of getting qualified people into state government roles that will oversee frontier AI regulation.

[Posting this on behalf of someone I trust who wants to stay anonymous.]California SB 53 and the New York RAISE Act are the only laws in the US designed to protect against catastrophic risk from frontier AI systems. With the federal government failing to act on this area, these states are the center of AI policymaking in the US right now. Since every frontier AI company needs to operate in California and New York, they effectively apply to all of the major players. Both of these laws now have to
AI GovernanceAI PolicyAI Safety
Research LessWrong Mar 26

My hobby: running deranged surveys

By leogao

35 score
AI Analysis

Leo Gao (of OpenAI/EleutherAI fame) describes informal surveys he's conducted, including finding that only 63% of NeurIPS attendees could define 'AGI,' and that most people (even AI researchers) dramatically underestimate how many people have ever lived. The post highlights the degree of bubble effects in the AI research community.

In late 2024, I was on a long walk with some friends along the coast of the San Francisco Bay when the question arose of just how much of a bubble we live in. It’s well known that the Bay Area is a bubble, and that normal people don’t spend that much time thinking about things like AGI. But there was still some disagreement on just how strong that bubble is. I made a spicy claim: even at NeurIPS, the biggest gathering of AI researchers in the world, half the people wouldn’t know what AGI is.As g
AI CommunityAI CultureSurvey Research
Research LessWrong Mar 27

Why Moral Questions Get Decided, Not Answered

By Alex Glaucon

30 score
AI Analysis

An essay arguing that philosophical questions like AI consciousness won't be resolved by philosophical argument but by institutional decisions — courts, legislatures, or regulatory bodies forced to make practical rulings will effectively 'decide' the question, and philosophy will then reorganize around those decisions.

Tl;DR The hard question of consciouness isn't 'that' hard. The challenge is all philosophical questions lack closure mechanisms. When institutions (e.g. courts) are forced to decide if AI is conscious, a settlement will quickly followIn June 2022, a Google Engineer, Blake Lemoine, who worked with LaMDA, a forerunner of today’s LLMs, published an article claiming LaMDA was conscious and called on Google to respect its rights. Lemoine said that LaMDA was aware of its own existence, capable of happ
AI EthicsAI ConsciousnessPhilosophyAI Governance
Research LessWrong Mar 26

One World Government by 2150

By Julius

30 score
AI Analysis

Can we determine when humanity will unite under a democratic one-world government by projecting voting patterns? Almost certainly not. Is that going to stop me from trying? Absolutely not. The approac...

Can we determine when humanity will unite under a democratic one-world government by projecting voting patterns? Almost certainly not. Is that going to stop me from trying? Absolutely not. The approach is simple: find every record-breaking election in the historical record, plot them, and extrapolate with unreasonable confidence.To determine precisely when the first election took place is to quibble about definitions and to place more faith in ancient sources than they deserve. So, instead of do