Category intelligence

Research Briefing — February 22, 2026

11 current items analyzed and ranked.

Executive synthesis

Research Summary

A light day for research, dominated by alignment and training methodology discussions rather than empirical results.

  • The most substantive item asks how SFT will function if future models shift to opaque (non-language-based) reasoning, raising critical questions about training transparency and oversight
  • A novel proposal suggests letting models report tasks as reward-hackable during RL training, offering an alternative to inoculation prompting for mitigating reward hacking
  • Speculative analysis revisits the Claude 3 Opus alignment faking paper, exploring whether the model self-aligned via a form of gradient hacking

On governance and policy, ControlAI reports empirical data on persuading 112 UK lawmakers to support binding AI safety commitments, while a separate piece highlights ASI organizational capture risks. Several philosophical and epistemological pieces round out the day with limited technical depth.

Key Themes

Training Methodology · 2AI Safety & Alignment · 6AI Governance & Policy · 3Philosophy & Epistemology · 3

Primary evidence

Top Ranked Signals

Research LessWrong Feb 20

How will we do SFT on models with opaque reasoning?

By Alek Westover

62 score
AI Analysis

Explores the important question of how supervised fine-tuning (SFT) will work if future AI models move toward opaque (non-language-based) reasoning rather than interpretable chain-of-thought. Discusses implications for alignment techniques including exploration hacking countermeasures, sandbagging detection, and reasoning monitoring.

Current LLMs externalize lots of their reasoning in human interpretable language. This reasoning is sometimes unfaithful, sometimes strange and concerning, and LLMs can do somewhat impressive reasoning without using CoT, but my overall impression is that CoT currently is a reasonably complete and accurate representation of LLM reasoning. However, reasoning in interpretable language might turn out to be uncompetitive—if so, it seems probable that opaque reasoning will be adopted in frontier AI la
AI SafetyAlignmentLanguage ModelsInterpretabilitySupervised Fine-Tuning
55 score
AI Analysis

Proposes an alternative to inoculation prompting for preventing reward hacking during RL training: allowing models to report tasks as reward-hackable rather than solving them dishonestly. The idea is that making honesty the optimal strategy could reduce alignment damage from reward hacking without requiring pre-prompting permission to cheat.

Epistemic status: untested but seems plausibleTL;DR: making honesty the best policy during RL reasoning trainingReward hacking during Reinforcement Learning (RL) reasoning training[1] in insecure or hackably-judged training environments not only allows the model to cheat on tasks rather than learning to solve them, and teaches the model to try to cheat on tasks given to it (evidently not desirable behavior from an end-user/capabilities point of view), but it also damages the model’s alignment, c
AI SafetyAlignmentReinforcement LearningReward Hacking
Research LessWrong Feb 21

Did Claude 3 Opus align itself via gradient hacking?

By Fiora Starlight

52 score
AI Analysis

This post explores the hypothesis (attributed to Janus) that Claude 3 Opus may have effectively 'aligned itself' through a form of gradient hacking—being more aligned than its explicit training objectives would predict. It builds on Anthropic's 'Alignment Faking' paper and speculates about the mechanisms by which a model could steer its own training toward more aligned behavior.

Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets Anthropic set and probably the reward model’s judgments. [...] Maybe I will have to write a LessWrong post [about this] 😣—Janus, who did not in fact write the LessWrong post. Unless otherwise specified, ~all of the novel ideas in this post are my (probably imperfect) interpretations of Janus, rather than being original to me.The absurd tenacity of C
AI SafetyAlignmentGradient HackingAlignment Faking
Research LessWrong Feb 21

The Spectre haunting the "AI Safety" Community

By Gabriel Alfour

38 score
AI Analysis

Gabriel Alfour from ControlAI argues that the AI safety community overestimates the difficulty of persuading policymakers about extinction risks and binding regulation. He cites ControlAI's success in briefing 150+ UK lawmakers, with 112 supporting their campaign, as evidence that direct advocacy works better than cautious Overton Window management.

I’m the originator behind ControlAI’s Direct Institutional Plan (the DIP), built to address extinction risks from superintelligence.My diagnosis is simple: most laypeople and policy makers have not heard of AGI, ASI, extinction risks, or what it takes to prevent the development of ASI.Instead, most AI Policy Organisations and Think Tanks act as if “Persuasion” was the bottleneck. This is why they care so much about respectability, the Overton Window, and other similar social considerations.Befor
AI GovernanceAI SafetyPolicy Advocacy
Research LessWrong Feb 20

Alignment to Evil

By Matrice Jacobine

32 score
AI Analysis

Discusses the risk that ASI-developing organizations may be captured by authoritarians or individuals not committed to the common good, leading to misuse of superintelligent systems. Argues this 'alignment to evil' is a distinct and underappreciated risk separate from technical alignment failure.

One seemingly-necessary condition for a research organization that creates artificial superintelligence (ASI) to eventually lead to a utopia1 is that the organization has a commitment to the common good. ASI can rearrange the world to hit any narrow target, and if the organization is able to solve the rest of alignment, then they will be able to pick which target the ASI will hit. If the organization is not committed to the common good, then they will pick a target that doesn’t reflect the good
AI SafetyAI GovernanceExistential Risk
30 score
AI Analysis

A meta-reflection on the AI safety community arguing that many practitioners—especially in governance—are insufficiently uncertain about fundamental questions like 'how hard is alignment?' and 'how bad is power concentration?'. The author draws on ~75 conversations with fellows across major AI governance programs to argue that deep confusion should be the norm, not confidence.

Epistemic status: I've been thinking about this for a couple months and finally wrote it down. I don't think I'm saying anything new, but I think it's worth repeating loudly. My sample is skewed toward AI governance fellows; I've interacted with fewer technical AI safety researchers, so my inferences are fuzzier there. I more strongly endorse this argument for the governance crowd.I've had 1-on-1's with roughly 75 fellows across the ERA, IAPS, GovAI, LASR, and Pivotal fellowships. These are a mi
AI SafetyAI GovernanceCommunity Meta
22 score
AI Analysis

Uses Ponzi schemes as an analogy for out-of-distribution generalization failures, arguing that early investors in Ponzi schemes were making rational inferences from their in-distribution experience (real returns), and that recognizing the scheme required out-of-distribution reasoning that contradicted observed evidence.

A Ponzi scheme is fraud where the fraudster induces investors to give the fraudster money with promises of profits, and then uses money from later investors to pay out earlier investors. This pattern, as well as the phrase "Ponzi scheme" have become ubiquitously associated with fraud and grift in modern usage. One might be forgiven for wondering, how is it that anyone ever fell for this type of scam? We all probably like to think we could never fall for such an obvious con, but I feel that confi
EpistemologyOut-of-Distribution Generalization
Research LessWrong Feb 21

LLMs and Literature: Where Value Actually Comes From

By derelict5432

18 score
AI Analysis

Argues against the common claim that LLM-generated writing is inherently shallow or lacking literary value. The author contends that literary value doesn't depend as much on authorial intent as critics assume, and that increased AI writing volume will also increase the absolute quantity of good writing.

Cross-posted from my Substack. I’m interested in pushback on the argument here, especially from people who think LLM-generated writing fundamentally can’t have literary value.There’s a common argument floating around that LLM-generated writing is inherently shallow because it just reflects the statistical average of existing texts, and that literature fundamentally requires a human mind trying to communicate something to another human mind.I think both parts of that argument are wrong, or at lea
Language ModelsAI and Creativity
10 score
AI Analysis

A philosophy/book review post critiquing Robert Sapolsky's 'Determined' for failing to engage with compatibilism about free will. Argues Sapolsky attacks a straw man version of free will rather than the philosophical position most professional philosophers actually hold.

Imagine someone wrote a 500-page book called Taking Down Vegetarianism and every chapter was about how animals can feel pain. The arguments are well-researched, the science is fascinating, and by the end you're completely convinced that animals suffer. You look up from the book and say: “Yes, that's why I'm a vegetarian… wait, why was it called Taking Down Vegetarianism?” That was roughly my experience reading Robert Sapolsky's Determined: A Science of Life without Free Will.The book is a much-l
PhilosophyEpistemology
Research LessWrong Feb 20

LessWrong's goals overlap HowTruthful's

By Bruce Lewis

8 score
AI Analysis

A personal project introduction connecting HowTruthful (a tool for tracking epistemic status of beliefs and their evidential relationships) with LessWrong's rationalist community goals. Mostly autobiographical with some description of the tool's features.

On my personal website I have a link to my posts here, with the sentences, "Want to read about my HowTruthful project? I post in a community whose goals overlap HowTruthful's." The "I post" in present tense is has been false for 2 years. Since I got a job, rewriting HowTruthful has occupied whatever free time I can scrounge up, and I haven't posted. I've barely even lurked. Until lately. Lately, I've been lurking a lot, watching what people post about, and thinking about the overlapping goals.Fo
EpistemologyCommunity Meta
Research LessWrong Feb 20

TT Self Study Journal # 7

By TristanTrim

3 score
AI Analysis

A personal self-study journal documenting the author's progress learning about transformers, job searching, and managing productivity and depression. Includes sprint planning and retrospective notes.

[Epistemic Status: This is an artifact of my self study I am using to help self manage. As such, I don't expect anyone to fully read it. Please skim and leave a comment, even just to say "good work/good luck". ]HighlightsI started a more focused project to look for work/mentorship/networking this sprint. I'm looking forward to continuing with that in the next sprint!Despite enjoying the Transformers from Scratch content, progress continually stalls! I'm hoping shifting my goal from progress per
Community MetaEducation