Category intelligence

Research Briefing — May 3, 2026

16 current items analyzed and ranked.

Executive synthesis

Research Summary

AI safety and alignment research dominates today's landscape, led by empirical work on monitor robustness and critiques of frontier-lab governance frameworks.

  • Fabien Roger's study measures Claude Opus 4.5's ability to generate adversarial inputs that fool narrow classifiers, directly stress-testing the viability of AI monitoring pipelines
  • A detailed critique argues OpenAI's Preparedness Framework v2 sets its "Critical" threshold for AI self-improvement dangerously late, firing only at 5x acceleration
  • Analysis of the Claude Mythos system card demonstrates that independent safety research techniques are being adopted in frontier deployments, validating the academic-to-industry pipeline

In evaluation and benchmarking, the Metaculus Summer 2026 FutureEval Bot Tournament ($50k) continues competitive assessment of frontier model forecasting. A novel 420-question benchmark on Nigerian ethnoveterinary knowledge highlights critical gaps in AI performance on non-Western domains. Additional contributions include a conceptual extension of mode collapse dynamics to human organizations, a preliminary policy sketch for government AI risk agility plans, and safety analysis of humanoid robots in light of Gemini Robotics ER-1.6 developments.

Key Themes

AI Safety & Control · 5AI Governance & Policy · 3AI Evaluation & Benchmarks · 3AI Existential Risk Communication · 2Psychology & Philosophy · 5

Primary evidence

Top Ranked Signals

Research LessWrong May 2

Measuring the ability of Opus 4.5 to fool narrow classifiers

By Fabien Roger

72 score
AI Analysis

Researchers measure Claude Opus 4.5's ability to generate adversarial attacks that fool prompted and fine-tuned classifiers (monitors) in a security-relevant context (BashBench). They find relatively low attack success rates, especially against chain-of-thought classifiers and fine-tuned Haiku 4.5, and present a methodology for evaluating future AI monitors.

We measure the ability of Opus 4.5 to fool prompted or fine-tuned classifiers trying to detect a narrow set of outcomes. We find that:The Opus 4.5 attacker gets a relatively low attack success rate on finding jailbreaks in BashBench, even when given some hints. Performance is especially low against a prompt Opus 4.5 classifier with a CoT and a fine-tuned Haiku 4.5 classifier. (There is almost no trusted monitor vs untrusted agent gap here, this mostly about whether LLMs are vulnerable against ja
AI SafetyAI ControlRed TeamingLanguage Models
Research LessWrong May 2

OpenAI's red line for AI self-improvement is fundamentally flawed

By Charbel-Raphaël

68 score
AI Analysis

A detailed critique of OpenAI's Preparedness Framework v2 'Critical' threshold for AI self-improvement, arguing it fires too late (5x acceleration vs Anthropic's 2x), is self-certified with zero external evaluators, and lacks measurable operational definitions. Proposes using METR's time horizon metric instead.

TL;DR. OpenAI's "Critical" threshold for AI self-improvement in the Preparedness Framework v2 has three structural problems:It fires too late. The lagging indicator, 5× generational acceleration sustained for several months, lets ~3 years of effective progress accumulate before triggering. Anthropic used a 2x threshold instead of a 5x.It's self-certified. Self-improvement is the only tracked category in the GPT-5.5 system card with zero external evaluators.It's not measurable. No operational def
AI SafetyAI GovernanceAI Self-ImprovementPreparedness Frameworks
62 score
AI Analysis

Analyzes the Claude Mythos system card to argue that independent technical AI safety research still matters, pointing out that recently-published techniques (Activation Verbalizers, emotion steering vectors) were directly used by Anthropic to detect misaligned behavior including cover-up actions.

When the Claude Mythos system card was released, I initially felt like we had entered a late stage of AI safety, where the number of parties that can make a real impact shrinks down to the handful of labs at the frontier, a few companies too critical to exclude from the conversation, and the governments of China and the US. Digging into the system card itself made me update significantly on this. Specifically, it reveals that amongst the techniques used in discovering misaligned behaviours were
AI SafetyAI AlignmentInterpretabilityResearch Strategy
45 score
AI Analysis

Metaculus announces its $50k Summer 2026 FutureEval Bot Tournament, continuing its series of AI forecasting benchmarks that pit frontier models against human forecasters. Provides a state-of-the-race update on how AI systems compare to professional human forecasters.

Summer Bot Tournament is StartingOver the last two years, Metaculus has been running a series of tournaments to benchmark AI's accuracy in predicting future events. These tournaments, now part of our broader FutureEval benchmark, pit frontier models, bot developers, and a human baseline against each other to collectively push the boundaries of forecasting performance. We are wrapping up the Spring Bot Tournament and are now prepping for the $50k Summer Bot Tournament!Joining the tournament is a
AI ForecastingAI EvaluationBenchmarks
Research LessWrong May 2

Evaluating different AI's on African livestck knowledge

By Fatika Umar Ibrahim

42 score
AI Analysis

A researcher builds a 420-question benchmark testing AI models on Nigerian ethnoveterinary practices, indigenous livestock breeds, and disease recognition. Llama 3.1 8B scored 43%, highlighting a significant knowledge gap in AI systems for African agricultural contexts.

I have been running evaluations on a niche that has almost zero attention in the AI safety world. Meta open source mode the llama 3.1 8b scored a 43% accuracy score on a 420 question benchmark I built covering ethnoveterinary practices, indigenous breed characteristics, disease recognition, and production systems specific to Nigeria.This evaluation is important because most other evals are ran on properly documented western specific data problems,. This project tests a domain where almost none o
AI EvaluationAI SafetyGlobal South AIBenchmark Development
Research LessWrong May 2

You Are Not Immune To Mode Collapse

By J Bostock

35 score
AI Analysis

An essay extending the concept of 'mode collapse' from AI image generation to human organizations and decision-making, arguing that feedback loops in grant-making, creative work, and specialization mirror the same dynamic where systems converge on narrow outputs.

“Mode collapse” is a few things. First it was an observation about how early image generating AIs often collapsed to producing just the modal output from their training distribution (something very common, like a house with a white picket fence and a tree in the garden). Then it was the observation that this effect seemed to occur extremely quickly when AIs were trained on AI-generated inputs. After that, it became the copium du jour of AI-is-hitting-a-wall folks for a while, who thought that th
Machine Learning ConceptsOrganizational TheoryAI Analogies
Research LessWrong May 2

AI Risk Agility Plans - v0.1

By Chris_Leong

32 score
AI Analysis

A brief policy proposal arguing governments need concrete 'agility plans' for AI governance rather than just stating intent to be agile. Suggests mechanisms including published plans, clear response timelines calibrated to capability progress rates, and guardrails against thrashing.

Epistemic status: Quick write-up to get the idea out there.“Plans are worthless, planning is everything” - Dwight D. EisenhowerMany governments say they want to be agile, I’ve seen this in many different policy documents, but words are cheap. If you want to be agile, you can’t just say you want to be agile, you need to design mechanisms to make it happen. I don’t want to be too prescriptive about the exact mechanisms here. Honestly, I suspect that it’s a bit like homework, a lot is lost if you j
AI GovernanceAI PolicyAI Risk
Research LessWrong May 1

Human-looking robots are a bad idea

By martinkunev

30 score
AI Analysis

Argues that humanoid robots with AI will amplify existing risks from AI chatbots (manipulation, parasocial relationships, privacy invasion) because embodiment triggers deeper human social instincts and creates new vectors for harm.

epistemic status: exploratory opinionated viewIt's not a coincidence that people have made cautionary tales about human-like robots. I want to share some thoughts on the issues stemming from human-AI interaction and argue that putting those AIs into human-looking robots would make the risks significantly worse.Current risks from advanced AIEvery now and then there is a scandal about some AI chatbots actively influencing people in dangerous ways. Those chatbots sometimes reinforce delusions, conv
AI SafetyRoboticsHuman-AI InteractionAI Ethics
Research LessWrong May 2

Understand why AI is a doom-risk in 39 captivating minutes

By KatjaGrace

28 score
AI Analysis

Katja Grace promotes an NPR podcast that presents her case for AI existential risk in 39 minutes, aimed at making the argument accessible to general audiences.

I’ve really wanted more good short accounts of why AI poses an existential risk. Working on one myself has been one of those incredibly high priorities I keep putting off. Meanwhile award-winning journalist Ben Bradford of NPR has made a podcast version of my case for AI x-risk that I am thrilled with! (Bonus within the 39 minutes: what Hamza Chaudhry of FLI thinks we should do about it—who I was delighted to later meet as a consequence!) If you or anyone you know could do with a quick and gripp
AI Existential RiskScience CommunicationAI Safety
Research LessWrong May 1

A Simulation of Social Groups Under A Gift Economy

By Mira Kennard

22 score
AI Analysis

Presents an agent-based simulation of gift economies inspired by Marshall Sahlins' 'Stone Age Economics,' modeling how reciprocal gift-giving dynamics emerge and function in pre-currency societies.

IntroductionI enjoy reading about people. Not individuals, but rather cultures, empires, kinship groups, etc. I'm fascinated by the emergent properties of multi-person systems. This post is born out of love for my favorite book: "Stone Age Economics" by Marshall Sahlins. In this text Sahlins elaborates on the economic life of hunter-gatherer and simple agricultural societies which exist today, extrapolating this into the past. He touches on many different points, but the one this post will focus
Computational Social ScienceEconomicsSimulation
Research LessWrong May 2

A new rationalist self-improvement book: the 12 Levers

By spencerg

20 score
AI Analysis

Spencer Greenberg announces a self-improvement book that systematically categorized ~500 techniques from 100+ popular self-help books and 20+ therapies into 12 high-level psychological strategies ('levers'), with evidence review for each.

I'm publishing a book that I think can fairly be described as a rationalist approach to self-improvement. Whereas many self-help books focus mainly on stories and what worked well for the author, our book takes a very different approach. My co-author, Jeremy Stevenson, and I (Spencer Greenberg) read over 100 of the most popular self-improvement books of all time and carefully reviewed more than 20 types of therapy in an attempt to answer the question: What are all of the most useful psychologica
Self-ImprovementPsychologyRationality
Research LessWrong May 2

Games that change your mind

By KatjaGrace

20 score
AI Analysis

Katja Grace shares a curated list of games that teach non-obvious real-world lessons through experiential play, including Dominion (don't invest for eternity), The Witness (nothing in the world tells you where to look), and various others.

Some things you might learn from games are pretty blatant: Trivial Pursuit might teach you trivia, MasterType might teach you about typing, Grand Theft Auto might teach you about driving or crime. But sometimes games teach people less obvious things—things that are more experiential or ineffable, things that you didn’t know you didn’t know, concepts that stick in your mind, deep things. Here’s my list of games and their interesting real-world updates, as experienced by me or my friends: Dominion
LearningGame DesignEpistemology