Explores the important question of how supervised fine-tuning (SFT) will work if future AI models move toward opaque (non-language-based) reasoning rather than interpretable chain-of-thought. Discusses implications for alignment techniques including exploration hacking countermeasures, sandbagging detection, and reasoning monitoring.
Category intelligence
Research Briefing — February 22, 2026
11 current items analyzed and ranked.
Executive synthesis
Research Summary
A light day for research, dominated by alignment and training methodology discussions rather than empirical results.
- The most substantive item asks how SFT will function if future models shift to opaque (non-language-based) reasoning, raising critical questions about training transparency and oversight
- A novel proposal suggests letting models report tasks as reward-hackable during RL training, offering an alternative to inoculation prompting for mitigating reward hacking
- Speculative analysis revisits the Claude 3 Opus alignment faking paper, exploring whether the model self-aligned via a form of gradient hacking
On governance and policy, ControlAI reports empirical data on persuading 112 UK lawmakers to support binding AI safety commitments, while a separate piece highlights ASI organizational capture risks. Several philosophical and epistemological pieces round out the day with limited technical depth.
Key Themes
Primary evidence
Top Ranked Signals
Reporting Tasks as Reward-Hackable: Better Than Inoculation Prompting?
By RogerDearnaley
Proposes an alternative to inoculation prompting for preventing reward hacking during RL training: allowing models to report tasks as reward-hackable rather than solving them dishonestly. The idea is that making honesty the optimal strategy could reduce alignment damage from reward hacking without requiring pre-prompting permission to cheat.
This post explores the hypothesis (attributed to Janus) that Claude 3 Opus may have effectively 'aligned itself' through a form of gradient hacking—being more aligned than its explicit training objectives would predict. It builds on Anthropic's 'Alignment Faking' paper and speculates about the mechanisms by which a model could steer its own training toward more aligned behavior.
Gabriel Alfour from ControlAI argues that the AI safety community overestimates the difficulty of persuading policymakers about extinction risks and binding regulation. He cites ControlAI's success in briefing 150+ UK lawmakers, with 112 supporting their campaign, as evidence that direct advocacy works better than cautious Overton Window management.
Discusses the risk that ASI-developing organizations may be captured by authoritarians or individuals not committed to the common good, leading to misuse of superintelligent systems. Argues this 'alignment to evil' is a distinct and underappreciated risk separate from technical alignment failure.
If you don't feel deeply confused about AGI risk, something's wrong
By Dave Banerjee
A meta-reflection on the AI safety community arguing that many practitioners—especially in governance—are insufficiently uncertain about fundamental questions like 'how hard is alignment?' and 'how bad is power concentration?'. The author draws on ~75 conversations with fellows across major AI governance programs to argue that deep confusion should be the norm, not confidence.
Ponzi schemes as a demonstration of out-of-distribution generalization
By TFD
Uses Ponzi schemes as an analogy for out-of-distribution generalization failures, arguing that early investors in Ponzi schemes were making rational inferences from their in-distribution experience (real returns), and that recognizing the scheme required out-of-distribution reasoning that contradicted observed evidence.
Argues against the common claim that LLM-generated writing is inherently shallow or lacking literary value. The author contends that literary value doesn't depend as much on authorial intent as critics assume, and that increased AI writing volume will also increase the absolute quantity of good writing.
A philosophy/book review post critiquing Robert Sapolsky's 'Determined' for failing to engage with compatibilism about free will. Argues Sapolsky attacks a straw man version of free will rather than the philosophical position most professional philosophers actually hold.
A personal project introduction connecting HowTruthful (a tool for tracking epistemic status of beliefs and their evidential relationships) with LessWrong's rationalist community goals. Mostly autobiographical with some description of the tool's features.
A personal self-study journal documenting the author's progress learning about transformers, job searching, and managing productivity and depression. Includes sprint planning and retrospective notes.