A Google DeepMind interpretability team update reporting that most safety-relevant properties of Gemini appear to arise from pretraining plus supervised fine-tuning rather than RL stages. They ran SFT on pretraining-only versions of Gemini 3.1 Pro and Gemini 3 Flash and found post-SFT models matched production models across safety benchmarks. This is an empirical, lab-internal finding with implications for where safety training effort should focus.
Category intelligence
Research Briefing — June 14, 2026
23 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's standout work clusters around safety/interpretability, governance, and a few niche applied items. Empirical results from a major lab lead the field.
Safety, Alignment & Interpretability
- Google DeepMind reports that most safety-relevant properties of Gemini arise from pretraining plus SFT rather than RLHF — a counterintuitive, actionable finding for alignment pipelines.
- An analysis of continual learning warns it may enable post-deployment goal and value drift, eroding current alignment guarantees.
- An AuditBench experiment shows a cheap Gemma 2-2B toxicity/EM judge gets adopted by investigator agents but fails to reduce alignment-audit costs (negative but useful result).
- Novel activation patching applied to a non-LLM DNA basecaller finds MLP dominance in early and late layers, extending interpretability tools to genomics.
Governance & Policy
- A US Commerce export-control directive suspended foreign access to Anthropic's Fable 5 and Mythos 5; covered via both Zvi's analysis and a primary-source statement.
- A strategy piece argues short timelines favor control, while long timelines favor infrastructure security, reframing timelines as non-uniform risk drivers.
- An AML-style compute-monitoring proposal aims to detect hidden or oversized training runs as a verification mechanism.
Applied ML & Discourse
- A peer-reviewed network with residual modules and attention targets jaw cyst image segmentation.
- A linkpost argues "AGI" has become nearly useless given highly jagged capabilities that decouple prior definitions.
Key Themes
Primary evidence
Top Ranked Signals
Part of a sequence analyzing how continual learning could reshape LLM agent safety, arguing it may allow post-deployment goal and value change and erodes the last-mover advantage of current safety interventions. It maps three pathways for value drift and three ways safety measures (pre-deployment evals, data filtering, control protocols) could weaken. A structured conceptual contribution to alignment thinking.
A cheap specialist judge gets used by agents but fails to reduce alignment audit costs
By burnssa
An empirical study giving AuditBench investigator agents a lightweight Gemma 2-2B toxicity/EM judge to test whether a cheap specialist tool reduces alignment audit costs. The agents reliably used the judge, but it only helped on quirks matching its training distribution and failed to lower total spend since the Sonnet driver dominated costs. A useful negative result for the alignment-auditing tooling community.
Zvi's commentary on a US Commerce Department export-control directive that suspended access to Anthropic's Fable 5 and Mythos 5 models for all foreign nationals, citing national security and a reported jailbreak. It analyzes whether this is targeted lawfare or broad national-security hawkishness and its implications for AI governance. The two affected models were released only days before this coverage date.
Short Timelines Favor Control, Long Timelines Favor Infrastructure Security
By Jannis
A strategy post arguing that AGI timelines do not uniformly change risk: longer timelines may reduce accidental misalignment but raise misuse and sabotage risks, so timeline length shifts which interventions have highest expected value. Drawing on the author's infrastructure-security background, it favors AI control under short timelines and infrastructure security under long ones. A conceptual contribution to safety prioritization.
US government directive to suspend access to Fable 5 and Mythos 5
By Capybasilisk
A primary-source repost of Anthropic's statement that a US government national-security export-control directive forced an abrupt shutdown of Fable 5 and Mythos 5 for all foreign nationals, reportedly over a jailbreak method. Anthropic notes the demonstrated vulnerabilities were minor and discoverable by other public models. It documents a notable regulatory intervention in AI deployment.
Exploration of a DNA Sequencing Basecaller using Activation Patching
By Madeleine L
An undergraduate mech-interp project applying activation patching to a non-LLM model, a DNA sequencing basecaller, finding MLP dominance in early and late layers, mid-layer attention activity, and concentrated activation in specific heads. It is exploratory interpretability on a biology pipeline component with potential relevance to pathogen surveillance.
A proposal to build an anti-money-laundering-style monitoring system that tracks AI infrastructure nodes to detect large or hidden training runs, supporting future compute-governance agreements like FLOP thresholds. It sketches an idea and references existing efforts like datacenter modeling and procurement MCPs. Early-stage concept rather than implemented research.
Multi-layer feature aggregation network with residual module and attention mechanism for jaw cyst image segmentation
By Unknown
A Nature Scientific Reports paper presenting a multi-layer feature aggregation network with residual modules and an attention mechanism for jaw cyst image segmentation. It targets improved medical image segmentation accuracy, though no abstract or details were provided in the source.
A linkpost arguing the term AGI has become nearly useless because AI capabilities are highly jagged, so different AGI definitions that once correlated now diverge sharply. It highlights how AI now does real economic work while remaining uneven across tasks. Mostly conceptual commentary on terminology.
A rebuttal to Ted Chiang's Atlantic essay on AI consciousness, agreeing LLMs are not conscious but disputing his claims that consciousness requires physical embodiment and that Anthropic implicitly denies Claude's consciousness. It is a reasoned commentary in the public AI consciousness debate.
Anthropic Is Taking AI Welfare Seriously. I’m Not Sure It Knows What It’s Measuring.
By Failfinder70
An opinion piece examining Anthropic's AI welfare efforts through the lens of Constitutional AI and RLAIF, arguing that current LLMs lack the persistent internal states needed for consciousness. It questions what exactly welfare measurements are capturing given the stateless, prompt-by-prompt nature of model inference. The post is commentary rather than original research.