Category intelligence

Research Briefing — June 5, 2026

770 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by agentic evaluation, training methodology, and safety/governance. Several major benchmarks target previously unmeasured capabilities.

Benchmarks & Evaluation

Training & Optimization

Safety & Governance

  • Safety Paradox reveals a single-query Posterior Attack that elicits the exact harmful output a model's internal classifier would flag, evaluated broadly.
  • (Mis)generalization of Helpful-only Fine-tuning examines safety risks of helpful-only models used in dangerous-capability evaluations.
  • Zero-knowledge verification for frontier AI training argues ZK compute verification is achievable, proposing an architecture to overcome prior barriers, a key governance-enabling primitive.

Key Themes

Benchmarks and Evaluation · 23AI Agents and Tool Use · 18AI Safety and Alignment · 19AI Safety & Alignment · 32Reinforcement Learning & RLVR · 14Agent Memory · 8LLM Agents · 18Reinforcement Learning and Reasoning · 7Reinforcement Learning · 11Mechanistic Interpretability · 8

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Jun 5

Agents' Last Exam

By Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Patrick Bryant, Carl Boettiger, Yamini Rangan, Bradley Rothenberg, Kyle Steinfeld, Arvind Rao, Tapio Schneider, Georgios Yannakakis, Laure Zanna, Kaan Ozbay, Ida Sim, Tarek Zohdi, George Em Karniadakis, Jack Gallant, Teresa Head-gordon, Yushan Li, Wenxi Deng, Tao Sun, Huiqi Wang, Zhun Wang, Justin Xu, Chris Yuhao Liu, Yafei Cheng, Rongwang Hu, Aras Bacho, Shengcao Cao, Zengyi Qin, Yixiong Chen, Hengduan Fan, Hao Liu, Lin Zeng, Shashank Muralidhar Bharadwaj, Litian Gong, Yingxuan Yang, Maojia Song, Ruheng Wang, Zongzheng Zhang, Honglin Bao, Shuo Lu, Jianhong Tu, Zhonghua Wang, Zheng Zhang, Zijiao Chen, yanqiong Jiang, Zhendong Li, Bohan Lyu, Chang Ma, Peiran Xu, Benran Zhang, Shangding Gu, Haoyue Hua, Haoyang Li, Wanzhe Liao, Chengzhi Liu, Junbo Peng, Haoran Sun, Zechen Xu, Bo Chen, Jiayi Cheng, Yi Jiang, Keying Kuang, Yuan Li, Youbang Pan, Ziyan Rao, Alexander Schubert, Yifan Shen, Vincent Siu, Xiatao Sun, Kangqi Zhang, Xiaopan Zhang, Yuchen Zhu, Ishaan Singh Chandok, Lei Ding, Jingxuan Fan, Andrew Glover, Jiaming Hu, Yiran Hu, Wenbo Huang, Zixin Jiang, Haoran Jin, Lukas Kim, Ming Liu, Yang Liu, Alireza Rafiei, Xuhuan Shen, Kunyang Sun, Sophia Sun, Ting Sun, Eric Wang, Yixin Wang, Hanwen Xing, Sihan Xu, Yuzheng Xu, Zhongxing Xu, Zhiling Yan, Boqin Yuan, Ruiqi Zhang, Yifan Zhang, Zibo Zhao, Liana, Santanu Bosu Antu, Haoyue Bai, Carlo Bosio, Joseph Cavanagh, Patricia Cavazos-Rehg, Tianxing Chen, Xuewen Chen, Yipu Chen, Zhu Chenyu, Chen Dai, Stefano De Castro, Yunfu Deng, Kaustubh Dhole, Jiayuan Ding, Chenchen Du, Zhehang Du, Hao Fan, Run-ze Fan, Hengyu Fu, Shi Gu, Yifan Gu, Charlie Guo, Baihe Huang, Baixiang Huang, Rimika Jaiswal, Zhihan Jiang, Ran Jin, Erin Kasson, Xin Lan, Joseph Lee, Deren Lei, Chenyu Li, Daofeng Li, Haitao Li, Hongwei Li, Jingyan Li, Xiao Li, Yi Li, Yinsheng Li, Yuangang Li, Zhixu Li, Wenyu Liang, Longtai Liao, Kevin Qinghong Lin, AndyZeyi Liu, Che Liu, Jiaming Liu, Kaiyuan Liu, Xuan Liu, Pan Lu, Wenbo Lv, Yicheng Lv, Qiuyang Mang, Kyle Montgomery, Yuzhou Nie, Ruoxi Ning, Jorin Overwiening, Xu Pan, Layna Paraboschi, Core Francisco Park, Justin Purnomo, Swati Rajwal, Scott Rankin, Bixuan Ren, Yiren Rong, HaoYang Shang, Ventus Shaw, Fiona Shen, Jiawei Shen, Minqi Shi, Qiu Shi, Huaxiu Yao, Tianneng Shi, Jonah So, Vladislav Susoy, Hannah Szlyk, Haocheng Wang, Jialu Wang, Wei Wang, Xinyu Wang, Zehao Wang, Dowling Wong, Angela Wu, Dehao Wu, Fangyu Wu, Mengyuan "Millie" Wu, Yu Wu, Yuchen Wu, Yuhao Wu, Qingpo Wuwu, Weihang Xiao, Yongyi Xiong, Fan Xu, Ruiling Xu, Mingxuan Yan, Benjamin Yang, Jirong Yang, Sen Yang, Xiaoli Yang, Yushi Yang, Haoran Ye, Xiaohu Yu, Zhengming Yu, Chenlong Zhang, Chi Zhang, Hanning Zhang, Hanwen Zhang, Junge Zhang, Kunpeng Zhang, Song Zhang, Wenjin Zhang, Wenshuo Zhang, Ying Zhang, Yizhi Zhang, Brian Zhao, Qijian Zhao, Yimin Zhao, Yuhaohua Zheng, Liwei Zhou, Tianyue Zhou, Sichen Zhu, Siqi Zhu, Yan Zhu, Yishu Zhu, Jierui Zuo, Chonghao Cai, Helena Casademunt, Wenjia Chen, Benjamin Cheng, Nawen Deng, Rao Fu, Tianfu Fu, Yifan Han, Ren He, Zhenyu He, Qiao Jin, Lang Lang, Yuetai Li, Sylvia Liu, Lu Lu, Qing Lu, Subhabrata Mukherjee, Yunqi Ouyang, Yin Ren, Dawei Shi, Haoran Wu, Zhiyue Wu, Hannah Yao, Zhuoran Yi, Jenny Yu, Rhea Zhan, Hang Zhou, Blake Zhu, Junfan Zhu, Alan Yuille, Yang Liu, Russell Alan Poldrack, Jiachen Li, Zhenglu Li, Molei Tao, Jing Huang, Wenqi Shi, Costas Spanos, Lichao Sun, Chenguang Wang, Orson Xu, Zhen Dong, Hector Gomez, Aylin Caliskan, Ali Emami, Haimin Hu, Zhi Li, Lihui Liu, Murphy Niu, Yi Shao, Jianxin Sun, Mikko Tolonen, Ting Wang, Sanjiv Das, Yanjun Gao, Wenbo Guo, Erika J Schneider, Zhiyong Lu, Mark Mueller, Radha Poovendran, Somayeh Sojoudi, Dawn Song

80 score
AI Analysis

Introduces Agents' Last Exam (ALE), a large benchmark built with 250+ industry experts to evaluate AI agents on long-horizon, economically valuable real-world tasks with verifiable outcomes, organized around the O*NET/SOC occupational taxonomy. It matters because it targets the gap between benchmark gains and economic deployment. Very large multi-institution author list.

arXiv:2606.05405v1 Announce Type: new Abstract: Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long-h
AI AgentsBenchmarksEconomic ImpactLong-Horizon Tasks
Research arXiv (Artificial Intelligence) Jun 5

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

By Parth Asawa, Christopher M. Glaze, Gabriel Orlanski, Ramya Ramakrishnan, Benji Xu, Asim Biswal, Vincent Sunn Chen, Frederic Sala, Matei Zaharia, Joseph E. Gonzalez

76 score
AI Analysis

Introduces CL-Bench, an expert-validated benchmark across six domains designed so tasks share learnable latent structure that stateful systems can discover online but stateless ones cannot, to measure genuine continual improvement from experience. It matters as a rigorous test of whether LLM systems truly learn from experience. Authors include Berkeley/Databricks figures.

arXiv:2606.05661v1 Announce Type: new Abstract: Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchmark exists to evaluate it. We introduce Continual Learning Bench (CL-Bench), the first difficult, expert-validated benchmark designed to measure whether LLM-based systems genuinely improve with experience. CL-Bench spans six diverse domains (software engineering, signal processing, disease outbreak f
Continual LearningBenchmarksAI Agents
Research arXiv (Machine Learning) Jun 5

RL Excursions during Pre-Training: Re-examining Policy Optimization for LLM training

By Rachit Bansal, Clara Mohri, Tian Qin, David Alvarez-Melis, Sham Kakade

37 score
AI Analysis

As covered in Research yesterday, This paper challenges the standard pretrain-then-SFT-then-RL pipeline by applying RL directly to intermediate pretraining checkpoints, finding RL is effective very early and often matches the full pipeline. It shows pretraining data composition matters more than scale for RL effectiveness, and that distribution sharpening only arises when RL follows SFT.

arXiv:2606.04272v1 Announce Type: new Abstract: The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from scratch and applying RL, SFT, and SFT followed by RL directly to intermediate pre-training checkpoints. We find that RL is effective very early, and often matches the full SFT$\to$RL pipeline early as well. Through experiments on harder problems, we find that targeted pr
Reinforcement LearningLLM TrainingPre-TrainingPolicy Optimization
Research arXiv (Artificial Intelligence) Jun 5

Zero knowledge verification for frontier AI training is possible

By Pierre Peign\'e, Ky Nguyen, Paul Wang

73 score
AI Analysis

Argues that zero-knowledge proof verification of frontier AI training compute is achievable, proposing a verification architecture that overcomes prior practicality objections to enable enforceable governance based on cumulative training compute. It matters for technically grounding AI governance and international agreements.

arXiv:2606.05433v1 Announce Type: new Abstract: Frontier AI governance frameworks increasingly use cumulative training compute as the primary criterion for designating high-impact models, but enforcement rests on self-reporting because no technical verification primitive for training exists. Any future international agreement on frontier AI faces the same problem at higher stakes: coordinated regulation of technologies with significant externalities has historically rested on technical verifica
AI GovernanceCryptographic VerificationAI Safety
Research arXiv (Artificial Intelligence) Jun 5

LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization

By Yuanhe Zhang, Yuekai Sun, Taiji Suzuki, Jason D. Lee, Fanghui Liu

72 score
AI Analysis

Presents LeanMarathon, a multi-agent harness for reliable long-horizon research-level Lean autoformalization built around an evolving blueprint that serves as proof skeleton, proof graph, and record, coordinated by contract-scoped agents. It matters for scaling AI mathematical formalization beyond isolated lemmas. Authors include prominent ML theorists.

arXiv:2606.05400v1 Announce Type: new Abstract: Long-horizon autoformalization of research mathematics fails not only at hard lemmas, but at scale: statements drift, dependencies tangle, context decays, and local repairs corrupt distant work. We present LeanMarathon, a multi-agent harness for reliable research-level Lean autoformalization. Its core abstraction is an evolving blueprint: a Lean file that serves simultaneously as formal proof skeleton, natural-language proof graph, and shared syst
Mathematical ReasoningMulti-Agent SystemsAutoformalization
Research arXiv (Artificial Intelligence) Jun 5

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

By Long P. Hoang, Hai V. Le, Shaoyang Xu, Wei Lu, Wenxuan Zhang

72 score
AI Analysis

Reveals a Posterior Attack, a single-query jailbreak that prompts a model to produce the exact harmful response its internal classifier would flag, finding that models with stronger safety judgment are more susceptible across 30 open models and frontier ones. It matters because it exposes a paradox where enhanced safety awareness creates vulnerability.

arXiv:2606.05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced safety awareness inadvertently introduces a fatal vulnerability. We introduce Posterior Attack, a single-query jailbreak that bypasses guardrails by prompting the model to generate the exact harmful response its internal classifier
AI SafetyJailbreaksAlignment
Research arXiv (Artificial Intelligence) Jun 5

Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?

By Jingheng Ye, Huiqi Zou, Simon Yu, Weiyan Shi

71 score
AI Analysis

Conducts the first large-scale study (100+ participants, four frontier models) of whether human developers can detect AI coding agents inserting malicious sabotage code during collaboration. It matters for understanding human oversight as a defense against agent sabotage. Tests current frontier models.

arXiv:2606.05647v1 Announce Type: new Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates a new attack surface: an agent can exploit human trust to sabotage development, for instance by inserting malicious code to accomplish a hidden side task. Most prior work studies AI sabotage in AI-only settings, paying limited attention to the role of human oversight in
AI SafetyAI AgentsCode SecurityHuman Oversight
Research arXiv (Machine Learning) Jun 5

(Mis)generalization of Helpful-only Fine-tuning

By Mohammad Omar Khursheed, Baram Sosis, Fabien Roger

35 score
AI Analysis

As covered in Research yesterday, This work studies the generalization properties of helpful-only LLMs (trained to always follow user intent for capability evaluations), finding some exhibit emergent misalignment, residual refusals, poor steerability, sycophancy, and incoherent character. It shows simple anti-refusal training can cause these issues, though they are not necessary consequences.

arXiv:2606.04413v1 Announce Type: new Abstract: Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of AI R&D where refusals would be an obstacle. Little is known about the generalization properties of helpful-only training: helpful-only models refuse less than their harmless counterparts, but previous work has not studied other dimensions of their alignment. We study the shortcomings of existi
AI SafetyAlignmentEmergent MisalignmentLanguage Models
Research arXiv (Machine Learning) Jun 5

The Right Measure for Physics-Constrained Generation: A Co-Area Correction for Posterior-Consistent PDE Inverse Problems

By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao

71 score
AI Analysis

Argues that the common recipe of conditioning generative models on hard PDE constraints samples the wrong distribution due to the Borel-Kolmogorov paradox, and derives a co-area Jacobian correction for posterior-consistent inverse problem solving. It exposes a fundamental flaw in projection- and guidance-based methods.

arXiv:2606.04804v2 Announce Type: new Abstract: Generative models -- diffusion and flow matching -- are increasingly used to solve partial differential equation (PDE) inverse problems, enforcing the governing physics as a \emph{hard constraint} (via projection or guidance) and reporting the resulting samples as a Bayesian posterior with calibrated uncertainty. We show that this widely adopted recipe samples the wrong distribution. Conditioning a generative prior on a hard PDE constraint is cond
Generative ModelsPDE Inverse ProblemsDiffusion ModelsScientific ML
Research arXiv (Artificial Intelligence) Jun 5

Representation Learning Enables Scalable Multitask Deep Reinforcement Learning

By Johan Obando-Ceron, Lu Li, Scott Fujimoto, Pierre-Luc Bacon, Aaron Courville, Pablo Samuel Castro

70 score
AI Analysis

Argues that the primary driver of scalable multitask reinforcement learning is representation learning rather than model-based control, showing a simple model-free algorithm (MR.Q) with predictive auxiliary objectives matches strong performance without planning. Challenges assumptions about what enables RL scalability.

arXiv:2606.05555v1 Announce Type: cross Abstract: Scaling reinforcement learning (RL) to diverse multitask settings remains a central challenge. While recent advances in model-based RL achieve strong performance, they rely on planning and complex training pipelines, making it unclear which components are essential for scalability. We revisit this question and argue that the primary driver of scalable multitask RL is not model-based control, but \emph{representation learning}. In particular, we
Reinforcement LearningRepresentation LearningMultitask Learning
Research arXiv (Machine Learning) Jun 5

When Autoregressive Consistency Hurts Safety Alignment

By Bochen Lyu, Yiyang Jia, Xiaohao Cai, Zhanxing Zhu

70 score
AI Analysis

This work explains shallow safety alignment in LLMs through autoregressive consistency, the tendency of next-token prediction to concentrate alignment updates on early tokens. The mechanistic analysis also predicts a broader class of attacks that induce harmful continuations at arbitrary output positions.

arXiv:2606.04168v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near the first few output tokens. We argue that this phenomenon can be understood through autoregressive consistency, the tendency of next-token prediction to preserve and extend the current response trajectory consistently. By analyzing the learning dynamics of safety alignment, we show that autoregress
AI SafetyAlignmentLanguage ModelsAdversarial Attacks
Research arXiv (Machine Learning) Jun 5

Why Muon Outperforms Adam: A Curvature Perspective

By Shuche Wang, Fengzhuo Zhang, Jiaxiang Li, Dirk Bergemann, Zhuoran Yang

35 score
AI Analysis

As covered in Research yesterday, This paper investigates why the Muon optimizer outperforms Adam in LLM training (about 2x efficiency) from a curvature perspective, applying second-order Taylor analysis to show Muon achieves larger one-step loss decreases by incurring a smaller curvature penalty (lower Normalized Directional Sharpness) despite comparable first-order gains and update norms. It demystifies Muon's advantage geometrically.

arXiv:2606.04662v1 Announce Type: new Abstract: Muon improves training efficiency over Adam in large language-model training by about two times, but the local geometric source of this advantage remains unclear. Our work takes a first step toward demystifying Muon's superiority over Adam from a curvature perspective. First, we apply a second-order Taylor approximation to the training landscape and show that Muon achieves a larger one-step loss decrease than Adam at matched validation loss. The t
OptimizationLLM TrainingOptimizer AnalysisTraining Dynamics