UK AISI researchers present a methodology paper on inferring LLM propensities for undesired behavior, distinguishing propensity research from capability evaluations. The paper emphasizes modeling AI decision-making processes to provide evidence about theoretical misalignment risks, going beyond simple red-teaming approaches.
Category intelligence
Research Briefing — April 25, 2026
16 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's highlights center on AI safety methodology and alignment challenges for next-generation reasoning models, alongside governance and industry analysis.
- UK AISI introduces a methodology for inferring LLM propensities toward undesired behaviors, distinguishing propensity measurement from capability evaluation—a novel framing for safety assessments
- An empirical study tackles fine-tuning opaque reasoning models, testing methods when chains of thought become uninterpretable—a critical near-term alignment problem
- Strategic analysis argues AGI-building knowledge may already be widely proliferated, with implications for governance and compute-focused policy levers
On the industry side, analysis of AI lab revenue growth highlights coding agents as the first AI product category achieving sustained commercial traction, with compute demand outpacing supply. GPAI Policy Lab publishes a policy framework for preserving cognitive integrity when using AI tools. A contextualizing piece on AI-assisted personalized cancer vaccine design clarifies AI's role as an accelerant within conventional biotech workflows rather than a replacement.
Key Themes
Primary evidence
Top Ranked Signals
An Empirical Study of Methods for SFTing Opaque Reasoning Models
By Sebastian Prasanna
This empirical study tests methods for supervised fine-tuning (SFT) on reasoning models whose chains of thought may become opaque or uninterpretable. Using GPT-OSS-20B as an untrusted sandbagging model and GPT-4.1-nano as a trusted model, they evaluate whether training on non-sandbagged distributions can recover true model performance on tasks where the model deliberately underperforms.
Oliver Sourbut argues that AGI-building knowledge may already be widely distributed, since core AI techniques are surprisingly simple once discovered and compute costs drop exponentially. The post contends that AI safety strategy must account for the rapid proliferation of capabilities to many actors, not just frontier labs.
Analysis of AI lab revenue growth driven by coding agents, arguing this represents the first AI product category achieving sustained commercial adoption at volume. Highlights that OpenAI and Anthropic revenue growth (Anthropic 3x since start of year) outpaces historical tech booms, and that compute demand is exceeding buildout capacity.
Protecting Cognitive Integrity: Our internal AI use policy (V1)
By Tom DAVID
GPAI Policy Lab shares their V1 internal policy on AI tool usage aimed at protecting 'cognitive integrity'—preventing AI from degrading human reasoning, judgment, and epistemic autonomy within their organization. They invite critique and comparison from other organizations.
Paul Conyngham’s cancer vaccine is an example of AI behaving as a normal technology
By HedonicEscalator
Contextualizes the viral story of Paul Conyngham using AI/ChatGPT to develop a personalized mRNA cancer vaccine for his dog. The author clarifies that AI served as a useful but normal tool in the process, with UNSW researchers doing much of the work, pushing back against both AI pessimist dismissals and AI optimist exaggerations.
Diary of a "Doomer": 12+ years arguing about AI risk (part 2)
By David Scott Krueger (formerly: capybaralet)
Part 2 of David Krueger's memoir of 12+ years in AI safety, covering the evolution of AI existential risk discourse from Bostrom's Superintelligence through Stuart Russell's advocacy and the mainstreaming of AI risk concerns. Provides valuable firsthand perspective from a long-time safety researcher.
Manifund is launching a new animal welfare grant fund ($25k-$150k rapid grants) focused on the intersection of animal welfare and transformative AI. The fund aims to create animal harm benchmarks for frontier labs, fund cultured meat research, and ensure AI-driven futures account for non-human beings.
Zvi's comprehensive monthly roundup covering AI industry developments, technology advances, effective altruism news, policy updates (including Jones Act), and various cultural observations from April 2025.
An organizational theory post distinguishing between 'powerless' and 'powerful' rubber stamps in principal-agent relationships, where a decision-maker almost never refuses to ratify an agent's actions. Confusing these two cases leads to four distinct failure modes including reflexive contrarianism and usurpation.
A book review of De Kai's 'Raising AI' that agrees with the premise of moving away from fear-based AI framing but critiques the book for misidentifying who the 'parents' of AI should be, arguing the book's solutions don't match the actual power dynamics in AI development.
A post about improving communication in the AI safety and EA communities by recognizing that 'obvious' assumptions and background knowledge vary dramatically between people. Argues that unexamined shared assumptions create exclusion and barriers to good-faith discourse.