Category intelligence

Research Briefing — May 9, 2026

14 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's most significant item is OpenAI's disclosure that chain-of-thought was accidentally graded during RL in several released models including GPT-5.4 Think, representing a concrete safety process failure with industry-wide implications.

  • ProgramBench faces validity critique arguing its near-zero frontier scores reflect impossible task design rather than capability gaps
  • Geoffrey Irving (UK AISI) frames alignment as choosing between adversarial vs. basin-of-attraction paradigms, with major strategic implications
  • "A benchmark is a sensor" proposes that benchmarks have sensitivity curves and are only informative within specific capability ranges
  • "Userland Alignment" identifies a neglected layer: building aligned systems through deployment harnesses rather than solely training-time interventions

On the applied side, AI capabilities are straining established coordinated disclosure and full disclosure security cultures simultaneously. Zvi Mowshowitz documents Claude Code performance bugs in April (now fixed) and practical agentic coding developments across tools.

Key Themes

AI Safety & Alignment · 4AI Benchmarks & Evaluation · 3AI Tools & Coding Agents · 2AI & Society · 3Community Meta · 3

Primary evidence

Top Ranked Signals

88 score
AI Analysis

Continuing our coverage from yesterday, OpenAI reports discovering that chain-of-thought (CoT) was accidentally graded during reinforcement learning in several released models (GPT-5.4 Thinking, GPT-5.1 Instant through GPT-5.4 Instant). This is significant because directly grading CoT can teach models to produce misleading reasoning traces, undermining a key safety monitoring mechanism. They describe their new automated detection system and the consequences of this accidental grading.

This is an unofficial automated linkpost. Monitoring our models’ chains of thought (CoT) has proven to be an effective way to detect and track model misalignment, both during RL training and deployment. While CoT monitoring has been useful for safety, we and many others in the industry believe CoT monitorability could be fragile. We would like to preserve and leverage CoT monitorability for as long as possible, and we recently introduced a suite of evaluations designed to measure it. Directly gr
AI SafetyAlignmentChain of ThoughtReinforcement LearningOpenAIModel Monitoring
Research LessWrong May 8

Is ProgramBench Impossible?

By frmsaul

52 score
AI Analysis

A critique of ProgramBench, a new coding benchmark where frontier models score near zero, arguing the benchmark is effectively impossible because its unit tests can capture obscure program behaviors that aren't discoverable through the 'clean room' black-box access provided to the coding agent.

ProgramBench is a new coding benchmark that all frontier models spectacularly fail. We’ve been on a quest for “hard benchmarks” for a while so it’s refreshing to see a benchmark where top models do badly. Unfortunately, ProgramBench has one big problem: it’s impossible!What is ProgramBench?ProgramBench tests if a model can recreate a program from a “clean room” environment. The model is given only a bit of documentation and black-box access to the program (all the programs are CLIs), then tasked
AI BenchmarksCode GenerationEvaluation Methods
Research LessWrong May 8

Bringing More Expertise to Bear on Alignment

By Edmund Lau

50 score
AI Analysis

Summarizes a keynote by Geoffrey Irving (UK AISI Chief Scientist) on the 'Adversaria vs Basinland' framing: alignment may be either an adversarial security problem or a navigational search problem, and we don't yet know which. Argues the field needs expertise from adjacent disciplines.

PreambleThe preamble is less useful for the typical AlignmentForum/LessWrong reader, who may want to skip to Adversaria vs Basinland section.On 28th of October 2025, Geoffrey Irving, Chief Scientist of the UK AI Security Institute, gave a keynote talk (slides) at the Alignment Conference. The conference was organised by the UK AISI and FAR.AI as part of the Alignment Project, which aims to bring experts from relevant fields to make progress on the alignment problem. TLDR:Adversaria vs Basinland.
AI AlignmentAI SafetyResearch StrategyAI Governance
Research LessWrong May 8

AI is Breaking Two Vulnerability Cultures

By jefftk

48 score
AI Analysis

Jeff Kaufman analyzes how AI acceleration is straining two established vulnerability disclosure cultures: 'coordinated disclosure' (tell maintainers privately, give them time) and Linux's 'open development' (fix bugs publicly but embargo security details). AI tools are collapsing the time window between patch publication and exploit development.

A week ago the Copy Fail vulnerability came out, and Hyunwoo Kim immediately realized that the fixes were insufficient, sharing a patch the same day. In doing this he followed standard procedure for Linux, especially within networking: share the security impact with a closed list of Linux security engineers, while fixing the bug quietly and efficiently in the open. His goal was that with only the raw fix public, the knowledge that a serious vulnerability existed could be "embargoed": the people
AI SecurityCybersecurityVulnerability DisclosureAI Acceleration
Research LessWrong May 8

Userland Alignment

By Josh H

45 score
AI Analysis

Proposes 'userland alignment' as a neglected approach: building aligned AI systems through carefully designed harnesses, system prompts, and deployment environments rather than solely through model training. Argues this is tractable for developers outside major labs.

Most discourse around AI alignment centers on model development and the labs that develop them. This is a reasonable place to focus given the centrality of model training to AI advancement. However, there are neglected opportunities to build defense-in-depth via aligned harnesses – and these opportunities might be tractable by interested developers and researchers who otherwise would struggle to have impact given the limited opportunities to influence lab practices.The behavior of an AI system i
AI AlignmentAI SafetyAI DeploymentDefense in Depth
Research LessWrong May 8

Claude Code, Codex and Agentic Coding #8

By Zvi

42 score
AI Analysis

Zvi Mowshowitz's eighth installment covering agentic coding developments, including Claude Code bugs that degraded performance in April (now fixed), Codex updates, and broader trends in AI-assisted coding. He notes the field has matured enough that standalone updates may fold back into weekly posts.

When I started this series, everyone was going crazy for coding agents. Now a lot more people are going crazy for coding agents, as well they should given how much better coding agents keep getting, but also Everybody Knows they are good and is focusing on actually using them. With the slower pace of news here it’s no longer clear that the waits associated with doing these updates on their own are worthwhile, so I’m going to fold these updates into the weekly again for now unless there’s a new m
Agentic CodingAI ToolsDeveloper Experience
Research LessWrong May 8

A benchmark is a sensor

By Håvard Tveit Ihle

40 score
AI Analysis

Presents a conceptual framework for understanding AI benchmarks as sensors with sensitivity curves: benchmarks are most informative within a specific capability range and lose discriminative power for models too weak or too strong. Discusses tradeoffs in benchmark design between sensitivity and range.

The simple mental pictureA simple mental picture we have for an AI capability benchmark is to think of it as a sensor with a certain sensitivity within a certain range of capabilities. The sensitivity of a benchmark, i.e. it's ability to distinguish the capability of different models, is given by a curve like this: The curve starts high (low sensitivity, high uncertainty), since for models with low capability all the tasks in the benchmark are too hard, and the benchmark can't distinguish betwee
AI BenchmarksEvaluation MethodsMeasurement Theory
25 score
AI Analysis

An essay arguing that chess is a poor analogy for predicting AI's cultural impact because people watch chess for the human drama between players, whereas they consume art (books, films) for the content itself. This distinction suggests AI-generated art may be more disruptive to creative fields than chess engines were to chess.

illustration by meIntroductionChess is sometimes used as an example of one of the first fields to benefit from the existence of artificial intelligence systems with superhuman performance. Both human performance and the enthusiasm surrounding human tournaments increased after the advent of such systems. In my opinion, these points are incorrectly cited as a basis for claims that affect broader fields.Different products generate interest because of their creators or their content. This distinctio
AI and SocietyAI Cultural Impact
Research LessWrong May 8

The Saturation View: some responses

By wdmacaskill

20 score
AI Analysis

Will MacAskill responds to critiques of his 'Saturation View' population axiology, which posits that the impersonal value of additional lives diminishes as more similar lives already exist. He addresses concerns about the view's novelty, its formal structure, and edge cases.

A couple of weeks ago, I published a draft of a new population axiology that I’ve been working on with Christian Tarsney. It got a lot of comments and pushback — thanks to everyone who engaged! They’ll feed into the more-polished academic-draft paper that Christian and I are working on.Here I’ll quickly respond to some of the most common or noteworthy responses. I’ll generally avoid stuff that is already covered in the draft.What’s the view? Isn’t this old hat?Very roughly, the Saturation view s
EthicsPopulation AxiologyEffective Altruism
12 score
AI Analysis

A post arguing that everyone in the EA/LW community who communicates publicly about high-impact ideas should receive media and communications training to improve clarity and reduce verbal noise. The author contends that poor oral communication reduces outreach effectiveness and credibility.

(This post is a first in a series of constructive criticism; reflections on improving the EA/LW ecosystem. High confidence for the general sentiment, uncertainty remains regarding the optimal mechanism for implementation. Post written by me, minor editing from Gemini.)SummaryEveryone who engages in oral communication about high-impact ideas should have training to improve their ability to do so (unless they are already proficient). Communications should be effective: clear, concise, linear, rati
Community BuildingScience Communication
Research LessWrong May 8

Please Be Serious

By Oliver Kuperman

10 score
AI Analysis

A post criticizing Eliezer Yudkowsky's decision to participate in a poorly-structured podcast debate with an anonymous 'AI lab director,' arguing it damaged his credibility and by extension the AI safety movement's reputation.

Recently, Eliezer Yudkowsky participated in a very flawed podcast of Doom Debates that reflected poorly on him, and, likely in the eyes of many, the entire AI safety movement. The premise of the debate was that Eliezer Yudkowsky was offered 10,000$ to debate an anonymous "AI lab director", and this director quickly made the debate into a mess by interrupting, yelling, and using profanity. Sure, Yudkowsky may have come across as sane in comparison, but his opponent did make one critical point dur
AI Safety MovementCommunity Building
Research LessWrong May 8

Write Cause You Have Something to Say

By Logan Riggs

5 score
AI Analysis

Advice about productive blogging, arguing that successful daily writers succeed because they have an 'overhang' of ideas to express, not merely because they write frequently. Offers tips on capturing ideas and building a writing pipeline.

The ones who are most successful at writeathons (Inkhaven, NaNoWriMo) are those with an overhang of things to say, usually in the form of:draft postsdaydreamsWhen Scott Alexander said:Whenever I see a new person who blogs every day, it's very rare that that never goes anywhere or they don't get good. That's like my best leading indicator for who's going to be a good blogger., it may seem you can just write every day, but that'd be Goodharting. There's something hidden in the writing process you
WritingProductivity