Category intelligence

Research Briefing — March 15, 2026

13 current items analyzed and ranked.

Executive synthesis

Research Summary

A thin day for research, led by a strong mechanistic interpretability result and a rigorous reanalysis of METR's developer productivity experiment.

Key Themes

Mechanistic Interpretability · 1AI Productivity / Economics · 1Recursive Self-Improvement / AI Capabilities · 2AI Governance and Public Opinion · 1AI Content Norms and Platform Policy · 1

Primary evidence

Top Ranked Signals

Research LessWrong Mar 14

Extracting Performant Algorithms Using Mechanistic Interpretability

By Ihor Kendiukhov

68 score
AI Analysis

Describes extracting interpretable algorithms from biological foundation models using mechanistic interpretability techniques, inspired by prior work finding evolutionary phylogenetic trees encoded in Evo 2's activations. Proposes that models trained on biological data (like single-cell data) encode performant algorithms that can be reverse-engineered.

A Prequel: The Tree of Life Inside a DNA Language ModelLast year, researchers at Goodfire AI took Evo 2, a genomic foundation model, and found, quite literally, the evolutionary tree of life inside. The phylogenetic relationships between thousands of species were encoded as a curved manifold in the model's internal activations, with geodesic distances along that manifold tracking actual evolutionary branch lengths. Bacteria that diverged hundreds of millions of years ago were far apart on the ma
Mechanistic InterpretabilityFoundation ModelsComputational BiologyAI for Science
62 score
AI Analysis

Performs a heterogeneity analysis of METR's late-2025 developer productivity experiment, finding that while the sample-wide AI speedup was ~6%, it was ~12% for tasks developers predicted would benefit from AI and up to ~25% for the most AI-proficient developer. Suggests the aggregate result masks meaningful variation.

Update: Fixed exponentiation of estimated parameters. Summary I use data from METR's recent developer productivity experiment to assess the possibility of heterogeneity in the effect of AI on time to complete a task. Relative to a sample-wide 6% speedup, I estimate a 12% speedup in tasks which were predicted by developers (prior to treatment assignment) to take substantially shorter with rather than without AI, and I estimate a 25% speedup for the developer in the study with the highest estimate
AI ProductivityDeveloper ToolsEmpirical AnalysisAI Economics
Research LessWrong Mar 14

Sparks of RSI?

By Nathan Helm-Burger

52 score
AI Analysis

Claims early signs of recursive self-improvement (RSI) are emerging in long-running AI agents that self-improve in loops with minimal prompting. Aggregates several Twitter posts as anecdotal evidence and predicts frontier labs will rapidly advance this capability.

Are your long-running agents self-improving in loops with minimal prompting? Mine sure are! I think we're seeing the first sparks of RSI here, folks. I'm expecting the frontier labs to scramble furiously to push this forward, finding and patching the meta-failure-modes. Thus, I expect next versions to be even better at this. Here's what some other people are saying/claiming: x.com/shreyasnsharma/status/2032567... x.com/varun_mathur/status/2032671842230501729 x.co/
Recursive Self-ImprovementAI SafetyAI CapabilitiesAlignment
Research LessWrong Mar 14

What concerns people about AI?

By spencerg

38 score
AI Analysis

Reports results from a US survey (October 2025) cataloging 16 distinct public concerns about AI, examining how worry levels vary by political orientation, gender, and AI knowledge. Provides a structured taxonomy of AI concerns and demographic breakdowns of who is most worried.

A lot of people are worried about AI. What are their worries? How worried are they? Are some demographics more worried than others? We ran a study to find out. In this article, we explain 16 concerns about AI that you might find it valuable to know about. We discuss, based on our data (collected in October 2025), how worried people in the US are about each concern.To whet your appetite, here are some questions that our study offers insights into. Can you predict what we found before we tell you
AI GovernancePublic OpinionAI SafetyAI Policy
Research LessWrong Mar 14

An AI skeptic's case for recursive self-improvement

By Harjas Sandhu

30 score
AI Analysis

A self-described AI skeptic lays out a step-by-step argument for why recursive self-improvement is plausible, starting from AI's code-writing ability through its application in ML research to potential feedback loops. Aimed at a general audience with a balanced presentation of reasons for doubt.

Note: this was originally written for a general audience. I'm posting it on Less Wrong because this community is much more informed about AI than the average person, and I expect that you have seen many of these arguments already—I would love to get your critiques / feedback.I’m not a huge believer in the intelligence explosion hypothesis—basically the idea that AI will become capable of self-improvement and thus speed up its own development.But I also don’t think the idea can be dismissed out o
Recursive Self-ImprovementAI CapabilitiesIntelligence Explosion
28 score
AI Analysis

Announces a major update to the LessWrong editor (now using Lexical framework with real-time autosave), and notably updates the site's LLM policy to allow inline AI-generated content blocks with clear attribution, demonstrated by having Claude Opus 4.6 write a section of the announcement post itself.

There's a new editor experience on LessWrong! A bunch of the editor page has been rearranged to make it much more WYSIWYG compared to published post pages. All of the settings live in panels that are hidden by default and can be opened up by clicking the relevant buttons on the side of the screen. We also adopted lexical as a new editor framework powering everything behind the scenes (we were previously using ckEditor).That scary arrow button in the top-left doesn't publish your post! It just op
AI PolicyAI-Generated ContentCommunity NormsLLM Integration
Research LessWrong Mar 14

FW26 Color Stats

By sarahconstantin

12 score
AI Analysis

Continuing a multi-season series, the author uses an automated pipeline with GPT-4o to classify colors in Fall/Winter 2026 fashion runway images from Vogue, finding an unusually dominant presence of black and burgundy as top non-neutral color this season.

Once again, I am coming out with stats on the colors in the latest fashion collections: this time, for the fall/winter 2026 season.Previous entries: SS26, FW25, SS25, FW24, SS24.Methodology RecapAs I’ve done before, I’m using my automated script for going through all the images hosted on Vogue Runway’s website for the current ready-to-wear collections, asking an LLM (currently GPT4o) to report all the colors in the outfit in each picture, and counting up the totals. So this means the “count” for
LLM ApplicationsData AnalysisComputer Vision
Research LessWrong Mar 14

Pragmatic approach to beliefs about consciousness

By Luck

10 score
AI Analysis

Proposes a pragmatic, instrumentalist approach to beliefs about consciousness: since phenomenal consciousness in others is unfalsifiable, one should strategically choose to believe entities are conscious when it produces net positive emotions. Suggests this as an optimization strategy for life quality.

Within the goal of maximizing one's own life quality, it is sometimes useful to believe that some entities are conscious. The property of phenomenal consciousness in others is unfalsifiable; therefore, whatever belief one has about the phenomenal consciousness of some other entity can't be false. This opens up the opportunity to strategically optimize this belief. There are pros and cons to believing that something has qualia. The belief that X has qualia allows one to associate with X, feel lov
ConsciousnessPhilosophy of MindEpistemology
Research LessWrong Mar 14

Optimal (And Ethical?) Methods To Find "Optimal Running"

By JenniferRM

8 score
AI Analysis

A personal blog post exploring the experience of using Google Search vs. Gemini to answer a question about optimal running, reflecting on the ethics of AI-generated search results and information quality. The author compares multiple methods of finding information and discusses moral qualms about Google's AI integration.

Epistemic Status: The central quote of this essay is just pure slop, of course. But argument screens off authority (or lack thereof), and I was genuinely curious about the object level answer, and I got the same rough answer from two methods (top hit vs trust Gemini), and the third method (read Gemini's links and think) had error bars and nuance that included the first two answers (but suggested ways to save some time every week). Editorial Status: I wrote this with the new editing tools of LW i
AI SearchInformation RetrievalAI Ethics
Research LessWrong Mar 14

Mini-Munich Succeeds Where KidZania Fails

By Novalis

5 score
AI Analysis

Compares two miniature city concepts for children—KidZania (scripted, commercial) and Mini-Munich (democratic, emergent)—arguing the latter succeeds because it gives children genuine agency and interconnected economic systems. Part of a larger exploration of whether miniature cities could replace traditional schooling.

This post is part of a larger exploration (not yet finished, but you can follow it at minicities.org) on whether a permanent miniature city could replace school. Tentatively, I think so, but the boundary between it and the adult world has to be deliberately porous, as I describe here.There are two well-known attempts to build miniature cities for children: Mini-Munich and KidZania.Both have streets, storefronts, jobs, and a local currency. But they are built on opposing assumptions about what ch
EducationSystems Design
Research LessWrong Mar 13

Sensing Physical Necessity: An Exercise In Naturalism

By Algon

5 score
AI Analysis

A phenomenological exercise exploring what physical necessity 'feels like' from a first-person perspective—trying to identify the sensory or cognitive signals that distinguish possible actions from impossible ones (e.g., walking through a column vs. walking around it).

Walking, I listen for sensations of physical necessity. It is surprisingly hard. Imagining myself jumping huge distances feels like imagining a normal step. Odd, that. Shouldn’t possible acts feel vivid, alive & real in a way fantasy isn’t?I make some progress by tracking what I feel are obstacles, like that medium sized column splitting the pavement, or that tree. They’re like gaps in my planned routes. Feels like I could circle around in two ways. On noticing this gap, I think wait a momen
PhenomenologyEpistemologyNaturalism
Research LessWrong Mar 14

'Staying with it' Done Wrong

By Selfmaker662

3 score
AI Analysis

A meditation reflection on how people commonly misapply the instruction to 'stay with' a feeling by locking onto a conceptual label rather than the living felt sense. The author draws on Gendlin's focusing technique to argue that feelings must be allowed to shapeshift rather than be pinned to a fixed concept.

I was meditating today and noticed quite some over-effort happening. So I did the diligent, spiritually respectable thing: I located it in the body — "pain in my forehead" — and decided to stay with it. I even felt a small glow of pride for remembering to find it somatically instead of getting lost on the mental level.What happened was quite disappointing: the headache locked into my attention, intensified steadily, and after a few minutes of escalating suffering I gave up and went to do somethi
MeditationPhenomenology