Category intelligence

Research Briefing — April 25, 2026

16 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's highlights center on AI safety methodology and alignment challenges for next-generation reasoning models, alongside governance and industry analysis.

  • UK AISI introduces a methodology for inferring LLM propensities toward undesired behaviors, distinguishing propensity measurement from capability evaluation—a novel framing for safety assessments
  • An empirical study tackles fine-tuning opaque reasoning models, testing methods when chains of thought become uninterpretable—a critical near-term alignment problem
  • Strategic analysis argues AGI-building knowledge may already be widely proliferated, with implications for governance and compute-focused policy levers

On the industry side, analysis of AI lab revenue growth highlights coding agents as the first AI product category achieving sustained commercial traction, with compute demand outpacing supply. GPAI Policy Lab publishes a policy framework for preserving cognitive integrity when using AI tools. A contextualizing piece on AI-assisted personalized cancer vaccine design clarifies AI's role as an accelerant within conventional biotech workflows rather than a replacement.

Key Themes

AI Safety & Alignment · 8AI Governance & Policy · 4AI Industry & Economics · 3AI Risk Communication & Community · 5

Primary evidence

Top Ranked Signals

Research LessWrong Apr 24

Methodology for inferring propensities of LLMs

By Olli Järviniemi

78 score
AI Analysis

UK AISI researchers present a methodology paper on inferring LLM propensities for undesired behavior, distinguishing propensity research from capability evaluations. The paper emphasizes modeling AI decision-making processes to provide evidence about theoretical misalignment risks, going beyond simple red-teaming approaches.

Our team at UK AISI has released a paper on inferring LLM propensities for undesired behaviour.I view this primarily as a methodology paper, and in this post I will talk about that:[1] First, I distinguish the aim of providing evidence on theoretical arguments regarding misalignment as separate from more red-teaming flavoured propensity research. Next, I discuss the methodological needs for providing such evidence, highlighting the need for modelling AIs’ decision-making. Finally, I give my pict
AI SafetyAlignmentLLM EvaluationMisalignmentAI Governance
Research LessWrong Apr 24

An Empirical Study of Methods for SFTing Opaque Reasoning Models

By Sebastian Prasanna

72 score
AI Analysis

This empirical study tests methods for supervised fine-tuning (SFT) on reasoning models whose chains of thought may become opaque or uninterpretable. Using GPT-OSS-20B as an untrusted sandbagging model and GPT-4.1-nano as a trusted model, they evaluate whether training on non-sandbagged distributions can recover true model performance on tasks where the model deliberately underperforms.

We open-source our code here.IntroductionCurrent reasoning models produce chains of thought that are largely human-readable, which makes supervised fine-tuning (SFT) on reasoning traces tractable: you can generate traces with a trusted model or by hand, and train on them directly. But it's not clear whether this will keep working. Future models may reason in ways that are hard to imitate—with chains of thought that use English in idiosyncratic ways, or even by reasoning in a continuous latent sp
AI SafetyAlignmentReasoning ModelsSupervised Fine-TuningSandbagging Detection
Research LessWrong Apr 24

Is the Cat Out of the Bag?: Who knows how to make AGI?

By Oliver Sourbut

42 score
AI Analysis

Oliver Sourbut argues that AGI-building knowledge may already be widely distributed, since core AI techniques are surprisingly simple once discovered and compute costs drop exponentially. The post contends that AI safety strategy must account for the rapid proliferation of capabilities to many actors, not just frontier labs.

Adapted from 2025-04-10 memo to AISII’ve previously made arguments like:Not long after it becomes possible for someone to make powerful artificial intelligence[1], it might become possible for practically anyone to make powerful AI.Compute gets exponentially cheaper by default.Knowledge proliferates (fast!) by default: AI techniques are typically simple and easy once discovered.What’s more, AGI-making know-how may be widespread already.Or, as Yudkowsky puts it[2],Moore’s Law of Mad Science: Ever
AI SafetyAI GovernanceAI Proliferation
Research LessWrong Apr 24

The World Can't Keep Up With AI Labs

By Lee.aao

38 score
AI Analysis

Analysis of AI lab revenue growth driven by coding agents, arguing this represents the first AI product category achieving sustained commercial adoption at volume. Highlights that OpenAI and Anthropic revenue growth (Anthropic 3x since start of year) outpaces historical tech booms, and that compute demand is exceeding buildout capacity.

Late last year a new AI psychosis kicked off. This time it was coding agents.People started saying this is a new era in programming, blah blah blah.*Karpathy tweet, late winter*A few months later, we’ve got more than just claims. We’ve got numbers. And they say something unusual is happening in the market.Coding agents are the first AI product people are paying for at volume and regularly. Because it directly speeds up their work. It’s too early to claim businesses are replacing whole processes
AI IndustryCoding AgentsAI EconomicsCompute Demand
35 score
AI Analysis

GPAI Policy Lab shares their V1 internal policy on AI tool usage aimed at protecting 'cognitive integrity'—preventing AI from degrading human reasoning, judgment, and epistemic autonomy within their organization. They invite critique and comparison from other organizations.

We (at GPAI Policy Lab) want to share our V1 policy as an invitation for pushback. Some of what motivates it is our extrapolations of AI capabilities, internal conversations about their effects on cognition, and some empirical evidence. I think the expected cost of being somewhat over-cautious here is lower than the cost of being under-cautious, and the topic deserves considerably more attention than it's currently getting. I'd love to see more orgs publish their own policies on this, both to co
AI SafetyAI GovernanceCognitive IntegrityOrganizational Policy
32 score
AI Analysis

Contextualizes the viral story of Paul Conyngham using AI/ChatGPT to develop a personalized mRNA cancer vaccine for his dog. The author clarifies that AI served as a useful but normal tool in the process, with UNSW researchers doing much of the work, pushing back against both AI pessimist dismissals and AI optimist exaggerations.

Submission note: I wrote this article last month, when this case was first reported. It has since been covered by Astral Codex Ten for April's linkpost, and was praised by RFK Jr. in a Senate hearing on Wednesday. (RFK Jr. was seemingly unaware that the AI-powered treatment he was referring to was an mRNA vaccine, a technology he has a history of opposing). This article aims to contextualize the role of AI in Conyngham's story.The Australian (archive link) recently reported that the entrepreneur
AI ApplicationsBiotechScience CommunicationAI Hype
Research LessWrong Apr 24

Diary of a "Doomer": 12+ years arguing about AI risk (part 2)

By David Scott Krueger (formerly: capybaralet)

28 score
AI Analysis

Part 2 of David Krueger's memoir of 12+ years in AI safety, covering the evolution of AI existential risk discourse from Bostrom's Superintelligence through Stuart Russell's advocacy and the mainstreaming of AI risk concerns. Provides valuable firsthand perspective from a long-time safety researcher.

Awareness and concern about the extinction risk posed by AI has been increasing the whole time I’ve been in the field. It feels like it’s finally going mainstream. But it’s also felt this way before……picking up where we left off in my previous post about how I got into AI and realized the field wasn’t thinking about x-risk…Nick Bostrom’s Superintelligence: Paths, Dangers, Strategies was widely criticized by the AI research community, but it did get the conversation started. None of the critiques
AI SafetyExistential RiskAI HistoryPersonal Narrative
Research LessWrong Apr 24

Manifund's Falcon Fund

By mabramov

25 score
AI Analysis

Manifund is launching a new animal welfare grant fund ($25k-$150k rapid grants) focused on the intersection of animal welfare and transformative AI. The fund aims to create animal harm benchmarks for frontier labs, fund cultured meat research, and ensure AI-driven futures account for non-human beings.

Manifund is launching a new animal welfare fund, led by regrantor Marcus Abramovitch. We make rapid (<1 week), early-stage ($25k–$150k) grants across animal welfare, with a particular interest in the intersection of animals and transformative AI.Reach out to marcus.s.abramovitch@gmail.com if you’d like to donate!Why AI x animals?Many EAs take seriously both the welfare of animals, and the possibility of short AI timelines. But EA funders currently consider these in isolation. AI safety grants
AI SafetyEffective AltruismAnimal Welfare
Research LessWrong Apr 24

Monthly Roundup #41: April 2025

By Zvi

20 score
AI Analysis

Zvi's comprehensive monthly roundup covering AI industry developments, technology advances, effective altruism news, policy updates (including Jones Act), and various cultural observations from April 2025.

AI continue to accelerate and dominate the schedule, which is why this is a bit late, but we do occasionally need to pay our respects to the Goddess of Everything Else. There’s cool or interesting things everywhere. Also maddenning things. But did you hear, for example, that they’re making some exceptions to the Jones Act? Table of Contents Bad News. Good Advice. Opportunity Knocks. Who Judges The Judges. Close Socrates. While I Cannot Condone This. Good News, Everyone. Violence Is Never The Ans
AI IndustryNews RoundupTechnology
Research LessWrong Apr 24

Rubber stamp errors

By jchan

15 score
AI Analysis

An organizational theory post distinguishing between 'powerless' and 'powerful' rubber stamps in principal-agent relationships, where a decision-maker almost never refuses to ratify an agent's actions. Confusing these two cases leads to four distinct failure modes including reflexive contrarianism and usurpation.

[Part of Organizational Cultures sequence]The "rubber stamp" is unduly maligned. When a Principal decision-maker is asked to ratify the actions of an Agent, but in practice never (or almost never) refuses, we call the Principal a rubber stamp. But this can mean one of two very different things:Powerless rubber stamp: The true power in fact rests with the Agent, even though in theory it "should" rest with the Principal.Powerful rubber stamp: The possibility that ratification may be refused incent
Organizational TheoryPrincipal-Agent ProblemsDecision Making
Research LessWrong Apr 23

Raising AI by Lowering Expectations

By Ramya

15 score
AI Analysis

A book review of De Kai's 'Raising AI' that agrees with the premise of moving away from fear-based AI framing but critiques the book for misidentifying who the 'parents' of AI should be, arguing the book's solutions don't match the actual power dynamics in AI development.

De Kai's Raising AI argues that fear-based framing in AI discourse is limiting us, and that we should think of AI as something we're raising rather than defending against. He's right about the framing but he's wrong about who the parents are - and the book inadvertently makes that case itself.In April, I took Bluedot Impact's Technical AI safety class. Throughout the readings, I kept noticing a pattern; AI safety researchers frequently discuss deceptive models, jailbreaks, and red teaming in lan
AI SafetyAI FramingBook Review
Research LessWrong Apr 24

Communicating with people who disagree on "obvious" things

By LawrenceC

14 score
AI Analysis

A post about improving communication in the AI safety and EA communities by recognizing that 'obvious' assumptions and background knowledge vary dramatically between people. Argues that unexamined shared assumptions create exclusion and barriers to good-faith discourse.

A commonly shared piece of wisdom in the LessWrong community is to say or do the obvious things. Normally, this is treated as an unambiguously good thing to do, for example, see Nate Soares’s “Obvious advice”.But I think there’s another genre of “obvious things” that requires more nuance: namely, the background assumptions that are so obvious to each of us or terms we hear so often they feel mundane. I think it’s uncontroversial to say that (“obviously”) what is obvious to you is often not obvio
Community BuildingAI Safety CommunityCommunication