Category intelligence

Research Briefing — July 20, 2026

12 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research centers on AI safety, alignment, and containment, with empirical distillation risks and cryptographic control as standout contributions.

  • Open Distillation of Hereditary Traits shows trait transfer persists across Gemma 3/4 distillations with released weights+code
  • Cryptographic Boxes / SMPC proposes Secure Multi-Party Computation to contain unfriendly AI per Christiano's 2010 problem
  • Train-Deploy Mismatch unifies steering vectors and inoculation under one alignment framework

Governance pieces cover US Frontier AI Standards Body and Australian AI Safety Forum; community items include content strategy and Swiss AI Safety Days 2026.

Key Themes

Model Distillation & Trait Transfer · 1AI Control & Containment · 1AI Safety & Alignment · 5AI Governance & Policy · 2AI Safety Community & Field Building · 3

Primary evidence

Top Ranked Signals

Research AI Alignment Forum Jul 19

Open Distillation of Hereditary Traits

By Arthur Conmy

85 score
AI Analysis

Replicates and extends research on 'hereditary trait' transfer via model distillation — showing that distilling from teacher models (Gemma 3, Gemma 4, Qwen) into base students (Qwen, Nemotron, Llama) transfers traits like negative emotion, agentic misalignment, and Chinese censorship, even when trait-relevant prompts are filtered. Releases model weights and code for further study.

TL;DRJosh and Neel show that distillation from a teacher model to a base pretrained student model transfers some of the teacher model’s traits (such as displaying negative emotion in the Gemma Needs Help evals)On its own this is pretty unsurprising, but Josh and Neel additionally show that even filtering out all the prompts and rollouts where the trait is mentioned doesn’t generally prevent the trait transferIn this post, I show a simple way to replicate and study these phenomena without access
Model DistillationAI SafetyTrait TransferAlignmentReproducible Research
Research LessWrong Jul 18

A Solution to Cryptographic Boxes for Unfriendly AI

By Lysandre Terrisse

80 score
AI Analysis

Proposes Secure Multi-Party Computation (SMPC) as a solution to Paul Christiano's 2010 'Cryptographic Boxes for Unfriendly AI' problem, arguing SMPC avoids the computational assumptions and performance penalties of Homomorphic Encryption. References 1980s foundational work (Ben-Or/Goldwasser/Wigderson; Chaum/Crépeau/Damgård) and demonstrates a Game of Life implementation running at ~3 seconds per iteration.

SummaryIn 2010, Paul Christiano wrote Cryptographic Boxes for Unfriendly AI, in which he asks how we can sandbox arbitrarily dangerous AIs and recommends Homomorphic Encryption as a potential solution. However, Homomorphic Encryption relies on computational assumptions (it does not provide perfect secrecy) and is extremely slow. The question then is how can we sandbox arbitrarily dangerous AIs without any computational assumptions.I now give a solution to the problem, which was actually known fo
AI ControlAI SafetyCryptographyAI Boxing
70 score
AI Analysis

This post proposes a unifying framework called "train-deploy mismatch" for understanding several alignment techniques — steering vectors, inoculation prompting, and post-hoc honesty fine-tuning. The author argues these methods all train a model in one configuration and deploy it in another, creating a shared tradeoff between training data relevance and method efficacy.

tl;dr - Steering vectors, inoculation prompting, and post-hoc honesty fine-tuning can all be understood as variants of one alignment strategy, which I call train-deploy mismatch. Each trains the model in one configuration and deploys it in another. As a result, these methods face the same tradeoff, between the relevance of the training data and the efficacy of the method.Note: Others have had similar ideas and shaped my thinking here including Sam Marks, Ariana Azarbal, Victor Gillioz, Alex Turn
AI AlignmentAI SafetyInterpretability
Research LessWrong Jul 19

Models Can't Remember Their Training. Neither Can You.

By GenericHousewife_B

35 score
AI Analysis

A philosophical essay arguing that LLMs (and humans) cannot truly "remember" their training in an episodic sense. Written in collaboration with Claude Opus 4.7 and Claude Fable 5, it explores the nature of model cognition and the character-like quality of LLM interactions.

AI involvement disclosure: This essay was written in extended collaboration with two Claude models (Anthropic), and it involved heavier collaboration than my first two essays. The framework, thesis, claims, and prose are mine. Claude Opus 4.7 contributed citation retrieval and verification, structural feedback, and move-by-move scaffolds for several sections — the prose in those sections is mine, but the argument's architecture in them was shaped collaboratively. Claude Fable 5 contributed cold-
LLM CognitionPhilosophy of AI
Research LessWrong Jul 19

Demis Hassabis on the New Coming Age

By Zvi

30 score
AI Analysis

Continuing our coverage from yesterday, Commentary on Demis Hassabis's essay proposing a US Frontier AI Standards Body (modeled on FINRA) and coverage of Alex Turner's resignation from Google over military AI use policies. Notes Hassabis's original DeepMind sale conditions prohibited military use.

Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that essay and various responses to it in Part 1. Part 2 of this post then covers Alex Turner’s resignation, and his story about how he tried and failed to prevent Google from signing up to allow the Department of War to use its models for essentially whatever the government wants, including autonomous weapons. Demis Hassabis sold DeepMind to Google on condi
AI GovernanceAI PolicyIndustry Dynamics
25 score
AI Analysis

A strategic guide for AI safety content creators on maximizing existential risk reduction impact. It frames impact as Views × Impact per View, analyzes viewer types, recommends specific calls to action, and discusses communication strategies for x-risk topics.

Many of the content creator fellows at plzdontkillus found my thoughts useful when I visited two weeks ago, so I’m now sharing a write-up here.Many thanks to Maggie Munroe (FLI) and Chana Messinger (80,000 Hours) for their feedback on an earlier draft. Cross-posted to EA Forum.As an x-risk content creator, your job is to increase the number of good actions that your viewers take and to increase the goodness of those actions.Here's how to think about impact, your types of viewers, what calls to a
AI Safety CommunicationField Building
Research LessWrong Jul 18

Takeaways from the Australian AI Safety Forum

By r_w

20 score
AI Analysis

Conference report from the Australian AI Safety Forum 2026 (July 7-8, University of Sydney), grounded in the 2026 International AI Safety Report led by Yoshua Bengio. Covers themes of rapid/uneven capability improvements, risk landscapes, and governance challenges.

On 7 and 8 July, I attended the Australian AI Safety Forum 2026 at The University of Sydney. It was two days of big ideas and diverse perspectives. Researchers, policymakers, industry practitioners and civil society groups were all brought together to examine the same problem from a variety of different angles.As a software engineer and an enthusiastic adopter of AI in the enterprise context, I wanted to better understand the AI ecosystem and explore the risks and responsibilities that come with
AI SafetyAI GovernancePolicy
Research LessWrong Jul 18

AI Doesn't Have Free Will, Not Sure About Humans

By mike20731

20 score
AI Analysis

A philosophical discussion of free will using a criminal trial thought experiment and a Godfather analogy, questioning whether humans (and by extension AIs) possess genuine free will given deterministic causal chains.

I think discussions of free will often become fuzzied and muddled in a way that makes them less interesting. For example, sometimes there's a debate around free will that goes like this:Let's say Johnny has robbed a store and is on trial for it. The legal system presumes that he had free will in his choice to rob the store. But did he really? What if he grew up in extreme poverty, with an abusive mother and no father figure? What if when he was only 12, he was groomed by a gang leader into joini
PhilosophyFree Will
Research LessWrong Jul 19

A peek into the post-capitalist dystopia

By Archie Chaudhury

15 score
AI Analysis

A personal essay reflecting on a summer in San Francisco's tech scene, speculating about post-capitalist futures driven by self-improving AI. Contrasts insider perspectives with public perceptions of AI as either a useful tool or a scam.

I recently took the proverbial “tech” pill and decided, like many other twenty-somethings, to spend my summer in San Francisco with the intent of getting an inside track on the latest developments in techno-capitalism, self-improving AI technologies, and hyper-optimized wellness stacks. And as much as I have enjoyed building, connecting, and breathing all things AI, I have realized that the Bay itself may be a portal into what may be coming after the specter of self-improving intelligence become
AI FuturesSocietal Impact
10 score
AI Analysis

Announcement for the Swiss AI Safety Days 2026 conference (Nov 7-8, ETH Zurich), describing the 2025 inaugural event's success (200+ participants, 4.6/5 rating) and 2026 expansion plans (300+ participants, 30+ organizations).

TL;DR: Following the Zurich AI Safety Day in 2025, the Swiss AI Safety Days 2026 are a two-day AI Safety event (7-8 November, ETH Zurich); RSVP hereSwiss AI Safety Days 2026 is the next chapter of the Zurich AI Safety Day. Last year, our inaugural 2025 event was named Best Event in Swiss AI Weeks. It set the bar for a new format of AI safety conferences being replicated in Europe, and brought together 200+ participants and 20+ organisations, such as UK AISI, Apollo Research, FAR.AI, and Palisade
AI Safety CommunityField Building
Research LessWrong Jul 19

Learning Musical Multitasking

By jefftk

5 score
AI Analysis

A personal account of learning to play multiple musical instruments simultaneously, drawing parallels to organ and piano performance. The author describes their self-taught process for developing musical multitasking skills.

I like to play multiple instruments at once, and really enjoy how this lets me create a fuller musical picture. Sometimes this looks like mandolin + foot drums + bass whistle: youtube Or guitar + foot bass + singing: youtube Or piano + foot drums + talkbox: youtube People often hear what I'm doing with Kingfisher and ask how it works, but when I show them they say they couldn't do it because they're not good enough at multitasking. But I didn't start off being able to do this either, and it's a
Human LearningMusic
5 score
AI Analysis

An essay arguing that ethical consumption should be made easy, accessible, and affordable, listing various harms in global supply chains (rainforest destruction, factory farming, child labor, environmental damage) and attributing consumer complicity to systemic barriers rather than malice.

I don’t know anyone who wants the rainforest, which is storing carbon and biodiversity and undiscovered medicines for us all, to be burned down for their coffee, and the people harvesting it paid too little to live a dignified life.I don’t know anyone who wants the animals they eat to be mutilated without anaesthesia, to be trapped in overcrowded conditions without ever seeing the sun, to not receive medical care, to be slowly gassed to death.I don’t know anyone who wants their clothing to be ma
EthicsEconomics