Category intelligence

Research Briefing — April 19, 2026

16 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by evaluation critique, AI control tooling, and alignment theory, with one empirical probe of a freshly released frontier model.

Benchmarks & Interpretability

AI Control & Safety

Research Meta

  • A 7-year post-mortem of early Assistive Multi-Armed Bandits/CIRL work and an essay on idea economics round out reflective commentary from established researchers

Key Themes

Interpretability & Evaluation · 2AI Safety & Control · 6Frontier Model Capabilities · 1Research Meta & Commentary · 4Off-topic / Personal · 4

Primary evidence

Top Ranked Signals

70 score
AI Analysis

Summarizes an EACL 2026 paper showing most Membership Inference Attack benchmarks detect distribution shifts rather than memorization (model-free baseline hits 98.6% AUC on VL-MIA-Flickr-2k), and introduces FiMMIA for multimodal MIA. Important critique of evaluation methodology.

Epistemic Status: I recently co-authored a paper on Membership Inference Attacks accepted at EACL 2026. More theoretical contributions — specifically the gradient attribution and the findings regarding the Hessian/positive-definite theories — are unpublished findings that I believe have some interest for AI Safety, Developmental Interpretability, and evaluation design. I am sharing them here for community feedback.Links: Paper | GitHub RepoTL;DR Many modern membership inference attacks (MIAs) ap
Membership InferenceBenchmarksMultimodalAI SafetyEvaluation
Research LessWrong Apr 17

Refactor Arena: A Control Setting for Software Engineering

By fastfedora

65 score
AI Analysis

Announces Refactor Arena, an open-source configurable AI control setting for studying whether agents inject vulnerabilities during code refactoring tasks. Extends the AI Control research agenda with a concrete, reusable benchmark.

Today we are announcing the release of Refactor Arena: a configurable, extensible control setting for exploring how agents might inject vulnerabilities into software applications while performing complex tasks such as refactoring code. In this post, we provide an overview of the setting and, at the end, the results of some initial evals done to validate the setting.However, the setting has been designed as a platform that can be configured and extended through YAML configuration files, and we en
AI ControlAI SafetyAgent EvaluationCode Generation
Research LessWrong Apr 18

Claude knows who you are

By Smaug123

55 score
AI Analysis

Replication of Kelsey Piper's observation that Claude Opus 4.7 can identify users from their writing style even when professing ignorance. Raises privacy and model-capability implications for frontier models.

Kelsey Piper noticed that Opus 4.7 is the first model which can identify her from her unpublished writing.I replicated the experiment myself, which is absolutely terrifying given that I am one of the most minor Internet personalities who has actually written stuff on the Internet.Claude professes not to know who I am, but reliably identifies me from my writing.Methodology: clear your custom instructions in claude.ai, and set your name to Unknown Visitor. Enter incognito chat mode with Claude. (A
Language ModelsPrivacyModel CapabilitiesAI Safety
Research LessWrong Apr 18

Latent Reasoning Sprint #4: PCA Analysis on CoDI

By Realmbird

45 score
AI Analysis

An interpretability experiment applying PCA and a tuned logit lens to the CoDI latent reasoning model (Llama 3.2 1B), finding that PC1 of hidden states correlates strongly with the end-of-chain-of-thought token. Offers novel mechanistic interpretability findings on latent reasoning representations.

In my previous post I found that activation steering worked with KV_cache and not with hidden state steering.So I decided to look at the PCA with methods such as logit lens and activation steeringQuick Summary:PC1 from hidden state activations strongly seems to correlate with the <|eocot|> or end of chain of thought token across all latent positions for the GSM8K datasetAdded Critiques of CodI near the endExperimental setupCoDI modelI use the publicly available CODI Llama 3.2 1B checkpoint
Mechanistic InterpretabilityLatent ReasoningLanguage Models
Research LessWrong Apr 18

Post-mortem'ing my earliest ML research paper, 7 years later

By LawrenceC

35 score
AI Analysis

A seven-year retrospective on the author's early ML paper on Assistive Multi-Armed Bandits/CIRL, reflecting on research direction choices. Useful meta-commentary on ML research processes.

Written quickly for the Inkhaven Residency.One of the things I like most about LessWrong yearly reviews is that they occur a full year after: for example, the reviews for posts written in 2019 happen at the end of 2020; we just reviewed the posts for 2024. As the LessWrong team writes:LessWrong has the goal of making intellectual progress on important problems. To make progress, you gotta examine your community's outputs not only when they're first published, but also once enough time has passed
Research MethodologyCIRLAssistance Games
Research LessWrong Apr 18

Don't Cut Yourself on the Jagged Frontier

By Against Moloch

30 score
AI Analysis

A thought experiment/dialogue exploring edge cases where even well-aligned superintelligence could cause harm through technologies humans can't verify. Contributes conceptually to safety discussions around verification and trust.

(With apologies to Sean Herrington, who deserves a better playwright than yours truly) A conversation with a friend on the bus to Bodega Bay today made me realize that there are some holes in my thinking about safety and superintelligence. I’ve assumed that superintelligence is by definition robustly better than humans at all the things, but there are some cases when that’s not the case. Without further ado, for your edification and discomfort, The Strawman Players present: A Disquieting Convers
AI SafetySuperintelligenceAlignment
Research LessWrong Apr 18

Down with the Old orthogonality thesis, up with the New

By Chris Santos-Lang

25 score
AI Analysis

An argument against Bostrom's instrumental convergence thesis, claiming empirical research falsifies the 'evil universe' version. Philosophical/theoretical contribution to AI risk discourse.

According to Wikipedia's article about Existential risk from artificial intelligence, one of the main worries of the x-risk school is that at least one of the instrumental goals upon which Superintelligence would converge happens to threaten our existence. Nick Bostrom raised this argument in 2012 (The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents) and offered up hoarding of resources as an example instrumental goal that would threaten our existence
AI SafetyAlignment TheoryOrthogonality Thesis
Research LessWrong Apr 18

Vladimir Putin's CEV is probably not that bad

By habryka

25 score
AI Analysis

Argues that even autocrats' coherent extrapolated volition wouldn't be catastrophic, pushing back on common CEV pessimism in alignment discussions. Contributes conceptually to alignment target discussions.

(Written quickly for Inkhaven, I hope someone someday makes a better case for this than I will here)Kelsey Piper on Twitter: me: it's not okay to hit your sister5yo: is it okay to kill Vladimir Putin?me: ...yes, if you were in a situation where it was somehow relevant it's okay to kill Vladimir Putin5yo: well, my sister is WORSE than Vladimir PutinNow, I do think Vladimir Putin is probably a pretty bad man all things considered. I personally am sympathetic to the current equilibria among major n
Alignment TheoryCEVAI Safety
Research LessWrong Apr 18

LLMs will soon disrupt algorithmic media feeds

By lsusr

20 score
AI Analysis

Prediction essay arguing LLM-powered blog curation startups will disrupt algorithmic media feeds by aligning with user rather than corporate values. Speculative commentary on LLM applications.

I predict that LLMs are about to disrupt algorithmic media feeds, and that this will start with a startup that curates blogs for you. Big Media is Misaligned If you look at a list of the world's top 10 websites, half of them are media websites. Of these 5 media behemoths, 4 (YouTube, Facebook, Instagram, X/Twitter) are misaligned [1] media feeds. By "misaligned media feed", I mean a website where the primary user interface is you go to the homepage and a giant machine learning algorithm shows yo
LLM ApplicationsMediaCommentary
Research LessWrong Apr 17

Idea Economics

By David Scott Krueger (formerly: capybaralet)

15 score
AI Analysis

Personal essay on the economics of research ideas and credit attribution, from David Krueger. Reflective but non-technical.

It happened again… I have an idea for a project. It’s a super cool project. I’d love for it to actually happen.The only problem is… I don’t have time to do it myself, and I want credit. But I don’t see a way to get it — well, not enough of it, at least. Having this experience a lot is one of the reasons I started Evitable: When I have a cool idea, I want to be able to execute it in house!As I was writing this, I realized I should share the example of the Statement on AI Risk, widely regarded as
Research CultureCommentary
Research LessWrong Apr 18

Book Review: The Unwritten Laws of Engineering

By Gordon Seidoh Worley

10 score
AI Analysis

A book review framing engineering self-help through Kegan's stages of adult development. Non-technical content with no AI research relevance.

There’s a genre of book that’s perennially popular. Some examples include:7 Habits of Highly Effective PeopleGetting Things DoneHow to Win Friends and Influence PeopleI’m Ok, You’re OkWhat these books have in common, aside from being self-help, is that they’re attempts to help people make the transition from the pre-rational, pre-systematic thought most of us have entering adulthood to the rational, systematic, modern, and self-authoring thought of Kegan Stage 4.This process is often plagued wit
Self-HelpCommentary
Research LessWrong Apr 18

If It's Worth Arguing, It's Worth Arguing With Whiteboards

By Drake Morrison

10 score
AI Analysis

Essay on productive disagreement and the value of whiteboards/visual tools in argumentation. General rationality content.

It's easy to disagree with people. You just say, "That's wrong" and decline to elaborate. But that's not very interesting. If you want to be making progress — instead of ragebaiting — it usually helps to find a way for your disagreement to be productive. Productive disagreements start in the place where you are already in agreement. The places where your models of reality overlap. Without any overlap at all, it gets quite difficult to successfully converse. You need to be speaking a language you
EpistemicsCommentary