Category intelligence

Research Briefing — June 19, 2026

544 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is anchored by a major frontier-model release and significant work in AI safety, evaluation rigor, and embodied learning.

Evaluation and reproducibility emerge as a strong theme:

Security, optimization, and robotics round out the list:

Key Themes

Language Models · 20Reinforcement Learning and Robotics · 12Vision-Language & Multimodal Models · 11Efficiency and Systems · 7Agentic AI and Multi-Agent Systems · 18LLM Agents and Coding · 9Diffusion & Generative Models · 26Evaluation and Benchmarks · 9AI Safety and Control · 4Benchmarking and Evaluation · 11

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Jun 19

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

By DeepSeek-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyu Hou, Chenhao Xu, Chenze Shao, Chong Ruan, Conner Sun, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Donghao Li, Dongjie Ji, Erhang Li, Fang Wei, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanting Chen, Guoai Cao, Guolai Meng, Guowei Li, Han Yu, Han Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoling Zhang, Haoming Luo, Haoran Wei, Haotian Yuan, Haowei Zhang, Haowen Luo, Haoyu Chen, Haozhe Ji, Hengqing Zhang, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J Yang, JQ Zhu, Jia Luo, Jia Song, Jia Yu, Jialiang Huang, Jialu Cai, Jian Liang, Jiangting Zhou, Jiasheng Ye, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jieyu Yang, Jin Chen, Jin Yan, Jingchang Chen, Jingli Zhou, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jingzi Zhou, Jinhua Zhu, Jiping Yu, Joseph Sun, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junmin Zheng, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Leyi Xia, Li Zhang, Liang Zhao, Lihua Guo, Lingxiao Luo, Linwang Ma, Linyan Zhu, Litong Wang, Liyu Cai, Liyue Zhang, Longhao Chen, MS Di, MY Xu, Max Mei, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Mingxu Zhou, Minmin Han, Ning Wang, Panpan Huang, Panpan Wang, Peixin Cong, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Qiwei Jiang, Rui Tian, Ruifan Xu, Ruijie Lu, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqian Chen, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, Ruyi Chen, SH Liu, Shanghao Lu, Shangmian Sun, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoheng Nie, Shaoqing Wu, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Shuying Yu, Songyang Zhou, Tao Ni, Tao Yun, Tian Jin, Tian Pei, Tian Ye, Tianle Lin, Tianran Ji, Tianyi Cui, Tianyuan Yue, Tingting Yu, Tun Wang, W Zhang, WL Xiao, Wangding Zeng, Wei An, Weilin Zhao, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjing Yao, Wenjun Gao, Wenkai Yang, Wenlve Huang, Wenqing Hou, Wentao Zhang, Wenting Ma, Xi Gao, Xiang He, Xiangwen Wang, Xianzu Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xiaowen Sun, Xiaoxiang Wang, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingchen Liu, Xingkai Yu, Xingyou Li, Xinyu Yang, Xinyu Zhang, Xu Chen, Xuanyu Wang, Xuecheng Su, Xueyin Chen, Xuheng Lin, Xuwei Fu, YC Yan, YQ Wang, YW Ma, Yanfeng Luo, Yang Zhang, Yanhong Xu, Yanru Ma, Yanwen Huang, Yao Li, Yao Li, Yao Xu, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, Yi Shao, Yi Yu, Yichao Zhang, Yifan Ding, Yifan Shi, Yijia Wu, Yiliang Xiong, Yiling Ma, Ying He, Ying Tang, Ying Zhou, Yingjia Luo, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiang Zhang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, YuKun Li, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuanhao Li, Yuduan Wang, Yuehan Yang, Yuer Xu, Yuhan Wu, Yuhao Meng, Yuheng Zou, Yukun Zha, Yunfan Xiong, Yupeng Chen, Yuping Lin, Yuqian Cao, Yuqian Wang, Yushun Zhang, Yuting Yan, Yutong Lin, Yuxian Gu, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuxuan Zhou, Yuyang Zhou, Yuzhen Huang, ZF Wu, Zehao Wang, Zehua Zhao, Zehui Ren, Zekai Zhang, Zhangli Sha, Zhe Fu, Zhe Ju, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zheren Gao, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhixuan Chen, Zhiyu Wu, Zhizhou Ren, Zhongyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihua Qu, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Ziyi Wan, Zizheng Pan, Zongqing Yao

88 score
AI Analysis

Technical report for the DeepSeek-V4 series, including a 1.6T-parameter (49B active) Pro model and 284B Flash model, both with million-token context via hybrid compressed-sparse attention, novel hyper-connections, and the Muon optimizer. Since these models reached GA in April 2026, this is the architectural report on an existing release rather than a new launch.

arXiv:2606.19348v1 Announce Type: cross Abstract: We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attenti
Language ModelsMixture-of-ExpertsLong-ContextEfficiencyModel Architecture
Research arXiv (Artificial Intelligence) Jun 19

Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots

By Khurram Javed, Joseph Modayil, Gloria Kennickell, Richard S. Sutton, John Carmack

78 score
AI Analysis

Presents Physical Atari, a robust hardware platform (Robotroller actuating an Atari controller plus a Devbox rendering frames/rewards) for studying real-time reinforcement learning on physical robots. Authored by Sutton, Carmack, and colleagues at Keen Technologies.

arXiv:2606.19357v1 Announce Type: cross Abstract: We built a robot called the Robotroller that actuates an Atari CX40+ controller and a device called the Atari Devbox that renders the game frame and the reward signal from the Arcade Learning Environment on a screen. The Robotroller and the Atari Devbox, together with an off-the-shelf camera and a desktop computer, constitute a system that can be used to study reinforcement learning algorithms in the physical world. We call the full system Physi
Reinforcement LearningRoboticsBenchmarks and PlatformsEmbodied AI
Research LessWrong Jun 18

GDM AI Control Roadmap

By Mary Phuong

76 score
AI Analysis

Google DeepMind published version 0.1 of an AI Control Roadmap describing internal guardrails to catch adversarial behavior by increasingly capable AI agents. It introduces a security-inspired threat taxonomy (TRAIT&R, building on MITRE ATT&CK) covering loss of control, work sabotage, and direct harm, plus defensive control invariants.

GDM has published an AI Control Roadmap! From the executive summary:We present the GDM AI Control Roadmap (v0.1) – our plan for implementing and adopting internal guardrails designed to catch potential adversarial behaviour by AI agents, even as they become increasingly harder to oversee and contain.We focus on system-level mitigations that limit the harm a misaligned AI system could cause. Specifically, this report provides:• Threat modelling: Taking inspiration from cybersecurity, we adopt a c
AI ControlAI SafetyThreat ModelingAdversarial Robustness
Research arXiv (Machine Learning) Jun 19

Optimal Deterministic Multicalibration and Omniprediction

By Georgy Noarov, Aaron Roth

72 score
AI Analysis

Resolves an open question by showing that deterministic predictors can achieve the minimax-optimal sample complexity for multicalibration and omniprediction, previously only attained by randomized predictors. This closes a gap in trustworthy ML theory.

arXiv:2606.20557v1 Announce Type: new Abstract: A model is multicalibrated on a collection of group weights $G$ if it is calibrated -- i.e. unbiased even conditional on its prediction -- not just overall, but also after reweighting contexts by each $g \in G$. It is a useful property for many downstream applications and is a basic desideratum of trustworthy machine learning. Before this work, all predictors known to attain the minimax-optimal $\widetilde O(\varepsilon^{-3})$ sample complexity ra
CalibrationTheoryTrustworthy MLFairness
Research arXiv (Artificial Intelligence) Jun 19

Human Universal Grasping

By Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto

70 score
AI Analysis

Presents HUG, a flow-matching model that generates diverse human grasps for any object from a single RGB-D image, trained on 1M-HUGs, a new egocentric dataset of 1M frames and 6,707 object instances collected via smart glasses. Argues human grasp data is the most natural source for generalizable robot grasping.

arXiv:2606.17054v1 Announce Type: cross Abstract: Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural source of robot grasping data is from humans, who pick up thousands of objects every day. We present HUG, a flow-matching model that generates diverse human grasps for any user-specified object in a single RGB-D image captured from a stereo camera. Using smart glasses, we first collect 1M-HUGs, an egocentric
RoboticsGraspingGenerative Models
Research arXiv (Artificial Intelligence) Jun 19

Playful Agentic Robot Learning

By Junyi Zhang, Jiaxin Ge, Hanjun Yoo, Letian Fu, Zihan Yang, Yaowei Liu, Raj Saravanan, Shaofeng Yin, Justin Yu, Dantong Niu, Zirui Wang, Roei Herzig, Ken Goldberg, Yutong Bai, David M. Chan, Ion Stoica, Angjoo Kanazawa, Jiahui Lei, Haiwen Feng, Trevor Darrell

70 score
AI Analysis

Introduces Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play (proposing tasks, executing code policies, diagnosing failures, distilling skills) before downstream tasks arrive. Strong Berkeley authorship (Goldberg, Darrell, Kanazawa, Stoica).

arXiv:2606.19419v1 Announce Type: cross Abstract: Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acquired only after explicit instructions. We study Playful Agentic Robot Learning, where an embodied coding agent uses self-directed play as a continual skill-learning stage before downstream tasks arrive. We introduce RATs, Robotics Agent Teams desig
RoboticsLLM AgentsSkill LearningEmbodied AI
Research arXiv (Machine Learning) Jun 19

VIMPO: Value-Implicit Policy Optimization for LLMs

By Zhewei Kang, Aosong Feng, Sergey Levine, Dawn Song, Xuandong Zhao

70 score
AI Analysis

VIMPO is a critic-free policy optimization method for LLMs that derives a policy-implied value function from the optimality conditions of KL-regularized RL, giving dense token-level credit assignment without a learned critic. It bridges the gap between simple group-relative methods like GRPO and unstable actor-critic approaches.

arXiv:2606.20008v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has become a central tool for improving the reasoning ability of large language models, but current methods face a trade-off between simplicity and credit assignment. Group-relative methods such as GRPO avoid training a critic, but typically assign a trajectory-level advantage to every token. Actor-critic methods provide denser learning signals, but require a learned value function with its own traini
Reinforcement LearningRLHFLanguage ModelsPolicy Optimization
Research arXiv (Machine Learning) Jun 19

FloatDoor: Platform-Triggered Backdoors in LLMs

By Nils Loose, Jonas Sander, Felix M\"achtle, Thomas Eisenbarth

70 score
AI Analysis

Introduces FloatDoor, the first input-independent backdoor attack against LLMs that triggers adversary-chosen behavior only when the model is served on a target deployment platform, exploiting non-associative floating-point arithmetic and divergent kernel implementations. Reveals a novel platform-dependent attack surface.

arXiv:2606.19535v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in sensitive settings such as software engineering, where their outputs directly shape downstream artifacts. Recent work has shown that an identical model can produce measurably different outputs depending on the deployment platform, a consequence of non-associative floating-point arithmetic and divergent kernel implementations. We study the security implications of this platform-dependent v
AI SecurityBackdoor AttacksLLMsSystems
Research arXiv (Computation and Language) Jun 19

Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias

By Justin D. Norman, Michael U. Rivera, D. Alex Hughes

70 score
AI Analysis

Conducts the largest systematic evaluation of LLM-as-a-Judge to date across 21 judges, three benchmarks, and ~541,000 judgments, showing that exact-match agreement overstates reliability and that chance-corrected kappa reveals universal deflation and unstable judge rankings. Audits agreement, consistency, and bias across the cohort including April 2026 frontier models.

arXiv:2606.19544v1 Announce Type: new Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for language models, but judge validation in practice relies on exact-match agreement, a metric that does not correct for chance and systematically overstates discriminative ability. We present the largest systematic evaluation of LLM-as-a-Judge to date: 21 judges from nine providers across MT-Bench, JudgeBench, and RewardBench, evaluated under three protocols (agreement, consistency, bias
LLM-as-a-JudgeEvaluationBiasLLMs
Research arXiv (Computer Vision) Jun 19

The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation

By Nicolas Dufour, Alexei A. Efros, Patrick P\'erez

70 score
AI Analysis

This paper treats FID as a random variable across training and sampling seeds, finding that retraining moves FID 3.2x more than resampling, exposing hidden randomness in generative model evaluation. It cautions against reporting single FID numbers from single models.

arXiv:2606.20536v1 Announce Type: new Abstract: The Frechet Inception Distance (FID) is the de facto arbiter of image generation, yet most papers report just a single number from a single trained model using a single sampling seed. How reproducible is that number if we retrain the model, or merely resample from it? In this paper, we treat FID as a random variable on a two-axis panel of training and generation seeds, and measure its variance directly on several hundred SiT networks trained on cl
Generative ModelsEvaluation MetricsReproducibility
Research arXiv (Artificial Intelligence) Jun 19

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

By Qian Zhao, Kunlong Chen, Changxin Tian, Zhonghui Jiang, Haitao Zhang, Chaofan Yu, Peijie Jiang, Mingliang Gong, Jia Liu, Ziqi Liu, Zhiqiang Zhang, Jun Zhou

69 score
AI Analysis

Identifies Shrinkage Bias as a fundamental limitation of non-uniform FP4 formats like E2M1 used in LLM pretraining, a systematic negative rounding error from geometric asymmetry that accumulates multiplicatively across layers and is amplified by the Random Hadamard Transform. Proposes a uniform-grid UFP4 recipe to explain and fix the observed training instability.

arXiv:2606.20381v1 Announce Type: new Abstract: FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes, including NVIDIA Blackwell/Rubin-class systems and AMD MI350-series GPUs, remain centered on E2M1 data elements. In this study, we identify a fundamental limitation of that choice: non-uniform formats such as E2M1 inherently suffer from Shrinkage Bias, a systematic negative rounding error caused by the geomet
Low-Precision TrainingEfficiencyHardware-Aware ML
Research arXiv (Computer Vision) Jun 19

World Engine: Towards the Era of Post-Training for Autonomous Driving

By Tianyu Li, Li Chen, Caojun Wang, Haochen Liu, Kashyap Chitta, Zhenjie Yang, Yuhang Lu, Naisheng Ye, Yihang Qiu, Yufei Wang, Luoxi Zou, Jiaxin Peng, Jin Pan, Zhaoyu Su, Andrei Bursuc, Shengbo Eben Li, Andreas Geiger, Peng Su, Hongyang Li

69 score
AI Analysis

World Engine is a generative framework that reconstructs high-fidelity interactive driving environments from real-world logs and extrapolates them into safety-critical long-tail scenarios for post-training autonomous driving policies. It targets the scarcity of rare high-stakes events in real datasets.

arXiv:2606.19836v1 Announce Type: cross Abstract: Autonomous vehicles must operate safely in the real world, where errors can have severe consequences. Although modern end-to-end driving policies excel in routine scenarios, their reliability is limited by the scarcity of safety-critical ``long-tail'' events in real driving datasets. These rare interactions define the practical safety boundary of the learned policy, yet they are difficult to collect at scale in the real world. Here we show that
Autonomous DrivingWorld ModelsGenerative Simulation