Category intelligence

Research Briefing — July 7, 2026

1139 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is anchored by Gemma 4, Google's open-weight multimodal family spanning dense and MoE variants (2B+) with a novel encoder-free design. Generative modeling advances with the first multiplayer interactive world model, conditioning on multiple agents' action streams to handle coupled physics in real time.

Safety, security, and alignment feature prominently:

Interpretability and training methods round out the top tier:

Key Themes

AI Safety and Alignment · 10LLM Agents · 57Reinforcement Learning · 31Language Models & Architectures · 6AI Safety, Attacks & Security · 13AI for Science · 15Reinforcement Learning & Distillation · 9Interpretability & Representation Geometry · 6Mechanistic Interpretability · 3AI Safety and Security · 10

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Jul 7

Gemma 4 Technical Report

By Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor C\u{a}rbune, Michelle Casbon, Mayank Chaturvedi, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Cl\'ement Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst, Jiaxian Guo, Cassidy Hardin, Yanzhang He, Steven M. Hernandez, Omri Homburger, L\'eonard Hussenot, Juyeong Ji, Armand Joulin, Aishwarya Kamath, Parnian Kassraie, Olivier Lacombe, Preethi Lahoti, Ga\"el Liu, Gus Martins, Luciano Martins, Tatiana Matejovicova, Ramona Merhej, Nikola Momchev, Sneha Mondal, Ryan Mullins, Sindhu Raghuram Panyam, Shreya Pathak, Sarah Perrin, Andr\'e Susano Pinto, Etienne Pot, Ang\'eline Pouget, Alexandre Ram\'e, Sabela Ramos, Douglas Reid, David Rim, Morgane Rivi\`ere, Karsten Roth, Louis Rouillard, Omar Sanseviero, Pier Giuseppe Sessa, Shane Settle, Danila Sinopalnikov, Sara Smoot, Piotr Stanczyk, Andreas Steiner, Lawrence Stewart, Ilya Tolstikhin, Michael Tschannen, Anton Tsitsulin, Nino Vieillard, Renjie Wu, Pingmei Xu, Haichuan Yang, Edouard Yvinec, Li Zhang, Joe Zou, Nicolas Aagnes, Abdelrahman Abdelhamed, Shivani Agrawal, Shubham Agrawal, Ibrahim Alabdulmohsin, Jean Baptiste Alayrac, Uri Alon, Chandramouli Amarnath, Ankesh Anand, Chrysovalantis Anastasiou, Setareh Ariafar, Fran\c{c}ois-Xavier Aubet, Kyriakos Axiotis, Federico Barbero, Joelle Barral, Alexei Bendebury, Urs Bergmann, Stanley Bileschi, Kat Black, Mathieu Blondel, Sebastian Borgeaud, Arthur Bra\v{z}inskas, Ryan Burnell, Robert Busa-Fekete, Mu Cai, Glenn Cameron, Charlotte Caucheteux, Garima Chadha, Jetha Chan, Aditya Chawla, Blake Jianhang Chen, Jesse Chen, Lin Chen, Xu Chen, Derek Cheng, Tzu-hsiang Chien, Nikolai Chinaev, Yi Chou, Zhaohui Chu, Benjamin Coleman, Pooja Consul, Sam Conway-Rahman, Scott Crowell, Dylan Cutler, Vivek Dani, Samira Daruki, Anil Das, Daniel Deutsch, Nishanth Dikkala, Li Ding, Qiuhan Ding, Shenil Dodhia, Konstantin Donhauser, Tulsee Doshi, Anca Dragan, Alex Druinsky, Sahil Dua, Zoltan Egyed, Danielle Eisenbud, Daniel Eppens, Cindy Fan, Bahare Fatemi, Yassir Fathullah, Vlad Feinberg, Milen Ferev, Takumi Fujimoto, Isaac Galatzer-Levy, Jo\~ao Gante, Simon Geisler, Soham Ghosal, Antonious M. Girgis, Alec Go, Alhaad Gokhale, Alex Grills, Yiming Gu, Pramod Gupta, Guru Guruganesh, Raia Hadsell, Hamza Harkous, Jitendra Harlalka, Demis Hassabis, Anja Hauth, Joe Heyward, Arian Hosseini, Chih-Yang Hsia, I-Hung Hsu, Xiaopeng Huang, Yangsibo Huang, Kevin Hui, Adrian Hutter, Te I, Fotis Iliopoulos, Advait Jain, Ganesh Jawahar, Ziwei Ji, Qilin Jin, Melvin Johnson, Kandarp Joshi, Arun Kandoor, Wang-Cheng Kang, Koray Kavukcuoglu, Mehran Kazemi, Kathleen Kenealy, Amr Khalifa, Phoebe Kirk, Suraj Kothawade, Vitaly Kovalev, Neel Kovelamudi, Adam Kraft, Ravin Kumar, Harish Kuppam, Justin Lannin, Chen-Yu Lee, Seungji Lee, Dmitry Lepikhin, Dongdong Li, Qiujia Li, Valentin Li\'evin, Ethan Lin, Ziqian Lin, Casper Liu, Tianlin Liu, Tianqi Liu, Xin Liu, Mayank Lunayach, Min Ma, Gagan Madan, Andrii Maksai, Eric Malmi, Michal Matuszak, Daniel McDuff, Gaurav Menghani, Daniil Mirylenka, Karolis Misiunas, Vedant Misra, Andreea Mitran, Kareem Mohamed, Maksim Mukha, Eric Noland, James O'Donnell, Kate Olszewska, Bernett Orlando, Wanqiong Pan, Rina Panigrahy, Unnati Parekh, Chunjong Park, Eric Paskie, Liqian Peng, Bryce Petrini, Slav Petrov, Jonas Pfeiffer, Bilal Piot, Martyna Plomecka, Siim Poder, Octavio Ponce, Arijit Pramanik, David Racz, Anish Rajan, Michelle Ramanovich, Anand Rao, Marvin Ritter, Vitor Rodrigues, Evan Rosen, Miko{\l}aj Rybi\'nski, Noveen Sachdeva, Micha\"el E. Sander, Rohit Sathyanarayana, Sagar Savla, Samuel Schmidgall, Tal Schuster, Benoit Seguin, Andrew Sellergren, Aliaksei Severyn, Izhak Shafran, Dhruv Shah, Yuan Shangguan, Ashish Shenoy, Pradeep Shenoy, Rakesh Shivanna, Pauline Sho, Lucas Spangher, Wojciech Stokowiec, Tim Strother, Yao Su, Yinghao Sun, Mukund Sundararajan, Andrea Tacchetti, Mor Hazan Taege, Pouya Tafti, Chetan Tekur, Rahul Thapa, Madeleine Traverse, Lenart Treven, Tao Tu, Chien Te Tung, Petar Veli\v{c}kovi\'c, Malini Pooni Venkat, Sagar Gubbi Venkatesh, Vidya Venkiteswaran, Francesco Visin, Alex Vitvitskyi, Kiran Vodrahalli, Weiyi Wang, Xin Wang, Tris Warkentin, Jan Wassenberg, John Wieting, Lechao Xiao, Hao Xu, Yuhui Xu, Fuzhao Xue, Arun Yadav, Jun Yan, Antoine Yang, Lin Yang, Ming-Hsuan Yang, Ziyu Ying, Jae Hyeon Yoo, Sajjad Zafar, Fred Zhang, Jiageng Zhang, Jianyi Zhang, Xiaofan Zhang, Chao Zhao, David Zhou, Chen Zou

80 score
AI Analysis

The Gemma 4 technical report introduces Google's new generation of open-weight natively multimodal models spanning dense and MoE architectures from 2.3B to 31B parameters, with improved vision/audio encoders, a unified encoder-free 12B model ingesting raw audio and image patches, and an integrated thinking mode. Gemma 4 was released in April 2026, so this documents an established model family.

arXiv:2607.02770v1 Announce Type: cross Abstract: We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture for our 12B model, which ingests raw audio and im
Language ModelsMultimodal ModelsOpen-Weight Models
Research arXiv (Artificial Intelligence) Jul 7

Multiplayer Interactive World Models with Representation Autoencoders

By Anthony Hu, V\'aclav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Am\'elie Royer, Manu Orsini, Alyx Liao, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Nor\'en, James Swingos, Jan H\"unermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz B\"ohle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick P\'erez

74 score
AI Analysis

This paper introduces the first multiplayer interactive world model for highly dynamic environments, conditioning on multiple agents' action streams to attribute scene changes to the correct player and stay coherent under arbitrary action combinations, demonstrated in Rocket League. The 5B-parameter latent diffusion model, trained on 10,000 hours of gameplay, generates real-time four-player matches at 20 fps on a single B200 GPU.

arXiv:2607.05352v1 Announce Type: cross Abstract: We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League
World ModelsDiffusion ModelsMulti-Agent SystemsGenerative AI
Research arXiv (Artificial Intelligence) Jul 7

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs

By Liyan Chen, Yael Tauman Kalai, Zoe Xi

73 score
AI Analysis

This theoretical AI-safety paper initiates the study of single-prover interactive proofs (specifically doubly-efficient ones) for verifying AI outputs, avoiding debate's assumptions that two provers are equally capable and one is truthful. It shows how to obtain verifiability guarantees without adversarial debate.

arXiv:2607.03561v1 Announce Type: new Abstract: As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions. A recent line of work focuses on verification via debate, a model of interactive proofs where two competing powerful provers, or AI models, debate each other to convince a weak verifier, or a human, of the correctness of their claim. However, debate assumes that the two AI models possess equal abi
AI SafetyInteractive ProofsScalable OversightTheory
Research LessWrong Jul 6

A global workspace in language models

By wesg

70 score
AI Analysis

The blog post for Anthropic's paper presenting evidence that language models like Claude have a small collection of verbalizable internal neural patterns functioning as a global workspace, analogous to consciously accessible processing, that can be described, controlled, and used for deliberate reasoning. It introduces techniques for identifying and accessing this space.

[This is the blog post for our new paper Verbalizable Representations Form a Global Workspace in Language ModelsReaders might also be interested in: the Public commentary, Github and Neuronpedia]As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that po
Mechanistic InterpretabilityLanguage ModelsAI Safety
69 score
AI Analysis

This paper audits whether using LLMs as evolutionary engines actually compounds discovery in scientific equation finding, and finds that parent-conditioned evolution is statistically indistinguishable from fresh independent sampling under matched budgets. It argues the loop effectively reduces to building a dictionary of candidate terms, then proposes set-level selection as a stronger alternative. It matters as a rigorous negative result challenging a popular LLM-for-science paradigm.

arXiv:2607.04108v1 Announce Type: new Abstract: Large language models are increasingly used as evolutionary engines for scientific discovery: generate candidates, select winners, feed them back as parents, and repeat. We audit whether this loop actually compounds discovery in scientific equation discovery, a setting where finite samples make structure underdetermined and interpolation easy. Under matched LLM-call budgets, parent-conditioned evolution is indistinguishable from fresh independent
Language ModelsScientific DiscoveryEvolutionary Methods
Research arXiv (Artificial Intelligence) Jul 7

LLM-as-a-Verifier: A General-Purpose Verification Framework

By Jacky Kwok, Shulu Li, Pranav Atreya, Yuejiang Liu, Yixing Jiang, Chelsea Finn, Marco Pavone, Ion Stoica, Azalia Mirhoseini

68 score
AI Analysis

LLM-as-a-Verifier proposes verification as a new scaling axis, computing continuous scores from the expectation over scoring-token logits rather than discrete judge outputs to provide fine-grained, training-free feedback for agentic tasks. The probabilistic formulation lets verification scale across multiple dimensions.

arXiv:2607.05391v1 Announce Type: new Abstract: Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In this work, we identify verification, the ability to determine the correctness of a solution, as a new scaling axis. To unlock this and demonstrate its effectiveness, we introduce LLM-as-a-Verifier, a general-purpose verification framework that provides fine-grained feedback for agentic tasks without requiring addi
Language ModelsVerificationTest-Time Compute
Research arXiv (Artificial Intelligence) Jul 7

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens

By Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks

68 score
AI Analysis

This interpretability study shows that frontier open-weight models (DeepSeek V3, Kimi K2) perform structured, legible multi-step computation over content-free filler tokens, with attention routing questions through the filler region and logit-lens readouts revealing intermediate reasoning. It matters for behavioral oversight because it demonstrates hidden reasoning is still decodable from internals.

arXiv:2607.03502v1 Announce Type: cross Abstract: Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). This is a limit case for behavioral oversight, where surface tokens carry no information about the underlying reasoning. But hidden from the output is not the same as hidden from us. On four task families (fact retrieval, parallel numeric composition, string manipulation
InterpretabilityAI SafetyLanguage ModelsReasoning
Research arXiv (Machine Learning) Jul 7

Untrusted Content Masking for Web Agents with Security Guarantees

By Kristina Nikoli\'c, Egor Zverev, Javier Rando, Matthew Jagielski, Edoardo Debenedetti, Florian Tram\`er

68 score
AI Analysis

Untrusted Content Masking extends provable prompt-injection defenses to web agents, which face a fundamental challenge because rendered pages intermingle trusted instructions with untrusted content, dissolving the trust boundary that guarantees rely on. It provides a masking mechanism to restore separation and offer security guarantees for web agents.

arXiv:2607.05277v1 Announce Type: cross Abstract: Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such as tool-use APIs, this separation arises naturally: agents can reason from interface definitions without ever processing untrusted content. Extending these guarantees to web agents faces a fundamental challenge: to perceive and interact with their environment, web age
AI SafetySecurityPrompt InjectionWeb Agents
Research arXiv (Computation and Language) Jul 7

Anchored Self-Play for Code Repair

By Caroline Choi, Zeyneb Kaya, Shirley Wu, Tengyu Ma, Tatsunori Hashimoto, Ludwig Schmidt

68 score
AI Analysis

Proposes generator-fixer self-play, where a single model is RL-trained to both generate bugs and fix them, creating an automatic curriculum as the fixer improves, and introduces BugSourceBench covering realistic bug sources. It scales code-repair supervision without limited human data.

arXiv:2607.03523v1 Announce Type: cross Abstract: Code repair is an important capability for language models (LMs): given a buggy program and unit tests, an LM must produce a fixed program that passes the tests. Because code repair data is limited, we aim to scale supervision by using an LM to generate bug--fix tasks. We propose __generator--fixer self-play__, in which a single model is trained with reinforcement learning to generate bugs and fix them. As the fixer improves, the generator adapt
Reinforcement LearningCode GenerationSelf-Play
Research LessWrong Jul 6

A Review of Anthropic's Global Workspace Paper

By Neel Nanda

68 score
AI Analysis

Neel Nanda's public review of Anthropic's global workspace paper, endorsing its evidence for a cognitive working-memory space in language models and the J-Lens technique for accessing it, and reporting replication of core claims on Qwen 3.6 27B plus preliminary evidence of abstract interpretative meta-tokens. It provides expert critical assessment and independent replication of an interpretability result.

The below is a public review Anthropic asked me to write for their new global workspace paper. I recommend at least skimming their paper first.TLDR:I think this is a fantastic paper - it presents compelling evidence for some kind of "cognitive space" in models, that is used as a "working memory" for intermediate variables during a forward pass, shows that J-Lens is a useful technique for accessing this space. I believe these key claims.I believe J-Lens will be a useful (but limited) tool in prac
Mechanistic InterpretabilityLanguage ModelsAI Safety
Research LessWrong Jul 6

Tie training can make DPO/RLHF-trained AIs generalize better

By Elliott Thornley

68 score
AI Analysis

This post summarizes an ICML paper arguing that DPO and RLHF cause models to latch onto every feature correlated with true value on the training distribution, including spurious ones, even in the infinite-data limit. The authors propose tie training, using pairs of equal-value actions with random or two-way labels, and show it reduces reliance on spurious features and improves out-of-distribution generalization.

This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training.TL;DROur theorems and experiments suggest that DPO and RLHF have an unwelcome consequence: they make AIs care about every feature of actions that correlates with true value on the training distribution.[1]That’s true even if the training set contains no misspecified preference data.And it’s true even in the infinite-data limit.
AlignmentPreference OptimizationRLHFGeneralization
Research arXiv (Artificial Intelligence) Jul 7

What Does a Discrete Diffusion Model Learn?

By Rodrigo Casado Noguerales, Bernhard Sch\"olkopf, Thomas Hofmann, Aran Raoufi

67 score
AI Analysis

This theoretical paper asks what a discrete diffusion model actually learns (a denoiser, score ratio, or bridge predictor), showing these are one object in different coordinates and that reading the network in the wrong coordinate changes the trained/sampled process. It rigorously derives the continuous-time Markov chain ELBO and proves an Oracle Distance theorem equating the negative ELBO exactly to data entropy plus path KL from the oracle reverse process.

arXiv:2607.05381v1 Announce Type: cross Abstract: What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of jump rates, these are one object in different coordinates, and reading a neural network in the wrong coordinate changes the process being trained and sampled. Starting with a rigorous derivation of the continuous-time Markov chain (CTMC) ELBO for any noising process, boundary terms included, we prove the \emph{Oracle Distance} th
Diffusion ModelsMachine Learning TheoryGenerative Modeling