Category intelligence

Research Briefing — July 19, 2026

8 current items analyzed and ranked.

Executive synthesis

Research Summary

Research centers on AI governance, interpretability, and alignment theory from practitioner and conceptual standpoints.

  • Red Line Framework ([61aae27f1df8]) proposes mechanism design for oversight in government AI contracts
  • Forbidden Technique ([3f7179b76bec]) critiques blanket bans on probe-based RLFR rewards via Goodfire Silico
  • Endogenous Alignment ([e0cadaaaded4], [8829f77d5a26]) contrasts internalized vs external value alignment
  • Payorian FairBot ([8775b2c5f442]) clarifies formal-agent equivalence in proof-based dilemmas
  • Remaining items are commentary, primers, or non-AI science with limited technical impact

Key Themes

Governance · 1AI Safety · 4Interpretability · 1Alignment · 4Decision Theory · 1Philosophy · 1Rationality · 1Biology · 1

Primary evidence

Top Ranked Signals

72 score
AI Analysis

A former Google DeepMind employee proposes a mechanism-design framework for setting red lines and oversight structures in government AI contracts, emphasizing robust language, minimal trust assumptions, and transparency via annual reporting. The piece draws on prior legal analysis of Anthropic's red lines and targets loophole-resistant governance of sensitive military and surveillance use cases.

My post on leaving Google DeepMind tells a story. In contrast, this Framework is a question of mechanism design and negotiation posture. I quite enjoyed optimizing this Framework against its organizational and practical constraints. The original considerations were:Good red lines: Rule out the questionable use cases (autonomous targeting without human control, untargeted profiling) while allowing trustworthy ones like missile defense. Avoid the weaknesses flagged in legal analysis of Anthropic’s
AI SafetyAlignmentPolicyGovernance
Research LessWrong Jul 17

The Most Forbidden Technique is not always forbidden

By Rauno Arike

68 score
AI Analysis

The post argues against a blanket ban on using model internals as training rewards, responding to Goodfire's Silico platform reproducing RLFR (probe-based RL). It reviews literature including The Obfuscation Atlas and clarifies conditions where training on internals is warranted.

A few days ago, Goodfire announced a private beta of Silico, their LLM training platform. As part of the announcement, they made a post describing Silico's reproduction of RLFR, a method developed by Goodfire that uses probes as reward signals for RL. Unsurprisingly, people on Twitter were quick to claim that "at long last, we have implemented the Most Forbidden Technique from the classic LessWrong post Don't Implement The Most Forbidden Technique".[1]As many have written before, blanket objecti
AI SafetyInterpretabilityAlignmentTraining Methods
Research LessWrong Jul 18

Endogenous Alignment

By Gordon Seidoh Worley

58 score
AI Analysis

The author introduces the distinction between exogenous alignment (external rewards and punishments) and endogenous alignment (internalized values) using childhood socialization as an analogy for AI alignment. The post argues adult-like agents may require endogenous rather than constant external control.

Starting when children are fairly young, usually around 1 year of age, we adults begin the work of aligning them to our values. We teach them to say “please”, not to hit, to ask for what they want instead of screaming, and much else. We do this primarily via exogenous methods, using a combination of punishments and rewards, that molds their behavior by encouraging good behaviors and discouraging bad ones.Such operant conditioning works because children have many instinctive behaviors that make t
AlignmentAI Safety
Research AI Alignment Forum Jul 18

Endogenous Alignment

By Gordon Seidoh Worley

56 score
AI Analysis

This is a cross-post of the endogenous alignment essay from the AI Alignment Forum, presenting the same analogy of exogenous versus endogenous value alignment using human socialization. It targets the alignment research audience specifically.

Starting when children are fairly young, usually around 1 year of age, we adults begin the work of aligning them to our values. We teach them to say “please”, not to hit, to ask for what they want instead of screaming, and much else. We do this primarily via exogenous methods, using a combination of punishments and rewards, that molds their behavior by encouraging good behaviors and discouraging bad ones.Such operant conditioning works because children have many instinctive behaviors that make t
AlignmentAI Safety
Research LessWrong Jul 18

My "Payorian FairBot" was just the original FairBot

By transhumanist_atom_understander

30 score
AI Analysis

The author notes that their proposed Payorian FairBot in a proof-based prisoner's dilemma tournament matches the original FairBot defined by MIRI. The post is a short correction within formal-agent and decision-theory circles.

MIRI's proof-based prisoner's dilemma tournament defined agents encoded as formulas of Peano arithmetic (PA) with one free variable. mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; font-weight: normal; font-size: 100%; font-size-adjust: none; letter-spacing: normal; border-collapse: collapse; word-wrap: normal; word-spacing: normal; white-space: nowrap; direction: ltr; padding: 1px 0; } mjx-container[jax="CHTML"][display="true"] { display:
Decision TheoryGame TheoryAgent Foundations
Research LessWrong Jul 18

The Coming of the Global Brain: A Review of "The God Test"

By Peter Kuhn

25 score
AI Analysis

A book review of Robert Wright's work on AI and evolutionary directionality, referencing Teilhard de Chardin's concept of a global brain. The post is commentary linking theology and AI speculation rather than original technical research.

I have published a brief review of Robert Wright's new book on AI and took the opportunity to discuss Teilhard de Chardin and the directionality of evolution:theanticompletionist.substack.com/p/the-coming-of-the-global-brain
AI CommentaryPhilosophy
Research LessWrong Jul 18

Map and Territory, Predictably Wrong

By manueldelrio

22 score
AI Analysis

A primer on rationality concepts from the LessWrong tradition, covering epistemic versus instrumental rationality, bias, and truth-seeking. It is introductory exposition rather than new research.

Crossposted (with small tweaks) from my Substack.What Do I Mean By “Rationality”?We distinguish epistemic rationality (making your beliefs more accurate) from instrumental rationality (succeeding at getting what you want).“rationality is about forming true beliefs[1] and making winning decisions”Main tools for this: probability theory (bayesianism), and decision theory.Feeling RationalBeing rational does not mean (contrary to traditional and popular expectations) being unemotional.Emotions, incl
RationalityEpistemics
Research LessWrong Jul 18

Nuances in the Workings of the Eye and Retina

By Hieronym

18 score
AI Analysis

An explanatory blog post surveys lesser-discussed evolutionary and structural details of the eye and retina, aimed at engineering-minded readers. It is biological exposition with no direct connection to AI systems or methodology.

Image edited from source at BritannicaThe eye (any eye) is a miracle of evolution, which has filled its structure at every scale with an apparent intentionality that is the hallmark of complicated systems under intensive selective pressure. It is very much possible to look at almost every design feature and state a compelling reason why it’s there, which makes it a rich place for the engineering-minded to find fun little design features.This post is meant to discuss topics that are less frequent
BiologyEvolution