Replicates and extends research on 'hereditary trait' transfer via model distillation — showing that distilling from teacher models (Gemma 3, Gemma 4, Qwen) into base students (Qwen, Nemotron, Llama) transfers traits like negative emotion, agentic misalignment, and Chinese censorship, even when trait-relevant prompts are filtered. Releases model weights and code for further study.
Category intelligence
Research Briefing — July 20, 2026
12 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's research centers on AI safety, alignment, and containment, with empirical distillation risks and cryptographic control as standout contributions.
- Open Distillation of Hereditary Traits shows trait transfer persists across Gemma 3/4 distillations with released weights+code
- Cryptographic Boxes / SMPC proposes Secure Multi-Party Computation to contain unfriendly AI per Christiano's 2010 problem
- Train-Deploy Mismatch unifies steering vectors and inoculation under one alignment framework
Governance pieces cover US Frontier AI Standards Body and Australian AI Safety Forum; community items include content strategy and Swiss AI Safety Days 2026.
Key Themes
Primary evidence
Top Ranked Signals
Proposes Secure Multi-Party Computation (SMPC) as a solution to Paul Christiano's 2010 'Cryptographic Boxes for Unfriendly AI' problem, arguing SMPC avoids the computational assumptions and performance penalties of Homomorphic Encryption. References 1980s foundational work (Ben-Or/Goldwasser/Wigderson; Chaum/Crépeau/Damgård) and demonstrates a Game of Life implementation running at ~3 seconds per iteration.
Many alignment techniques work by training one model and deploying another
By cloud
This post proposes a unifying framework called "train-deploy mismatch" for understanding several alignment techniques — steering vectors, inoculation prompting, and post-hoc honesty fine-tuning. The author argues these methods all train a model in one configuration and deploy it in another, creating a shared tradeoff between training data relevance and method efficacy.
Models Can't Remember Their Training. Neither Can You.
By GenericHousewife_B
A philosophical essay arguing that LLMs (and humans) cannot truly "remember" their training in an episodic sense. Written in collaboration with Claude Opus 4.7 and Claude Fable 5, it explores the nature of model cognition and the character-like quality of LLM interactions.
Continuing our coverage from yesterday, Commentary on Demis Hassabis's essay proposing a US Frontier AI Standards Body (modeled on FINRA) and coverage of Alex Turner's resignation from Google over military AI use policies. Notes Hassabis's original DeepMind sale conditions prohibited military use.
Stop Chasing Views: How to Reduce x-Risk as an AI Safety Content Creator
By Luc Brinkman
A strategic guide for AI safety content creators on maximizing existential risk reduction impact. It frames impact as Views × Impact per View, analyzes viewer types, recommends specific calls to action, and discusses communication strategies for x-risk topics.
Conference report from the Australian AI Safety Forum 2026 (July 7-8, University of Sydney), grounded in the 2026 International AI Safety Report led by Yoshua Bengio. Covers themes of rapid/uneven capability improvements, risk landscapes, and governance challenges.
A philosophical discussion of free will using a criminal trial thought experiment and a Godfather analogy, questioning whether humans (and by extension AIs) possess genuine free will given deterministic causal chains.
A personal essay reflecting on a summer in San Francisco's tech scene, speculating about post-capitalist futures driven by self-improving AI. Contrasts insider perspectives with public perceptions of AI as either a useful tool or a scam.
Save the date: Swiss AI Safety Days 2026 (7-8 November, ETH Zurich)
By andrejfsantos
Announcement for the Swiss AI Safety Days 2026 conference (Nov 7-8, ETH Zurich), describing the 2025 inaugural event's success (200+ participants, 4.6/5 rating) and 2026 expansion plans (300+ participants, 30+ organizations).
A personal account of learning to play multiple musical instruments simultaneously, drawing parallels to organ and piano performance. The author describes their self-taught process for developing musical multitasking skills.
Ethical consumption needs to be easy to identify, accessible and affordable, especially in a worsening economy. And we need to rethink what we should be optimising our economy for
By Portia
An essay arguing that ethical consumption should be made easy, accessible, and affordable, listing various harms in global supply chains (rainforest destruction, factory farming, child labor, environmental damage) and attributing consumer complicity to systemic barriers rather than malice.