Category intelligence

Research Briefing — July 29, 2026

120 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research highlights major advancements in open frontier architecture, state-based agent execution, and memory paradigms that eliminate inference compute. Key breakthroughs span open mega-scale MoEs, program-state computer interaction, and automated scientific code discovery.

Frontier Architectures & Memory Paradigms

  • Kimi K3 (Moonshot AI): Releases a 2.8T parameter Mixture-of-Experts (MoE) architecture with 104B active parameters and a 1M-token context window, setting a new open SOTA for massive-scale multimodal reasoning.
  • Persistent Solution Memory: Proves that pairing a frozen 12B model with a persistent, verified memory store yields 100% accuracy at 0 generation tokens for deterministic tasks, demonstrating a cost-effective alternative to continuous inference and re-computation.

Agentic Control & Reasoning Reliability

Multimodal & Generative Systems

Embodied AI & Scientific Discovery

Safety & Oversight Frameworks

Key Themes

AI Safety, Alignment & Governance · 10Agents & Automation · 12Agentic Frameworks & Evaluation · 9Robotics & World Models · 14Model Distillation & Efficiency · 9Multimodal & Vision-Language Models · 13Efficient Inference & Attention Mechanisms · 4Robotics & Vision-Language-Action Models · 5AI for Healthcare & Specialized Domains · 5

Primary evidence

Top Ranked Signals

Research Hugging Face Papers Jul 28

Kimi K3: Open Frontier Intelligence

By Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, Yuxuan Cao, Ziwei Chai, Y. Charles, H. S. Che, Guanduo Chen, Guangyu Chen, Guanzheng Chen, Huarong Chen, Jia Chen, Jianlong Chen, Jun Chen, Kexin Chen, Peng Chen, Ruijue Chen, Wentao Chen, Xin Chen, Yang Chen, Yanru Chen, Yifei Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Dazhi Cheng, Yean Cheng, Jialei Cui, Jingbing Cui, Anqi Dai, Jiaqi Deng, Hao Ding, Rui Ding, Shaofeng Ding, Mengfan Dong, Mengnan Dong, Yuhao Dong, Yuxin Dong, Angang Du, Chenzhuang Du, Dikang Du, Jusen Du, Yulun Du, Yu Fan, Jing Feng, Qiulin Feng, Yichen Feng, Kelin Fu, Qiang Fu, Fuxuan Gao, Hongcheng Gao, Jingyue Gao, Tong Gao, Weijia Gao, Shangyi Geng, Jie Gong, Linhu Gong, Shengao Gong, Xiaochen Gong, Qizheng Gu, Yicheng Gu, Shuhao Guan, Haiqing Guo, Shiqi Guo, Xiang Guo, Zhengyan Guo, Beixi Hao, Wenxin Hao, Xiaoru Hao, Dailan He, Haotian He, Lehan He, Qi He, Weiran He, Xinran He, Xinyi He, Yibo He, Yunjia He, Chao Hong, Tiange Hong, Hao Hu, Jiaxi Hu, Ruikun Hu, Weiming Hu, Yangyang Hu, Zhenxing Hu, Liang Hua, Jinbin Huang, Ke Huang, Ruiyuan Huang, Siying Huang, Weixiao Huang, Yan Huang, Zhengjie Huang, Zhiqi Huang, Yulong Hui, Chaobo Jia, Yutong Jiang, Zhejun Jiang, Zuoyou Jiang, Wenyi Jin, Xinyi Jin, Yu Jing, Huanjun Kong, Guokun Lai, Aidi Li, Cheng Li, Chengyuan Li, Cong Li, Fang Li, Guanyu Li, Haoyang Li, Jia Li, Junxiong Li, Lei Li, Letian Li, Lincan Li, Weihong Li, Wentao Li, Xintong Li, Yang Li, Yishen Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zhaoxi Li, Zheming Li, Zhengxiao Li, Zhiyuan Li, Jiawei Lin, Xiaohan Lin, Yibo Lin, Zichao Lin, Ziyan Lin, Bill Liu, Boxiao Liu, Chuan Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yipeng Liu, Zhengying Liu, Zhiheng Liu, Enzhe Lu, Haoyu Lu, Linqiang Lu, Tingzhan Lu, Zhiyuan Lu, Aotian Luo, G. Luo, Junyu Luo, Yifan Luo, B. Lyu, Wenzhou Lyu, Shaoguang Mao, Yuan Mei, Xin Men, Minqing Ni, Yixuan Niu, Siyuan Pan, Shujun Peng, Zhangyang Qi, Ruoyu Qin, ZeChao Qin, Zeyu Qin, Haiquan Qiu, Jianxin Qiu, Jiezhong Qiu, Bowen Qu, Yuhao Qu, Zeyu Shang, Youbo Shao, Han Shen, Jincheng Shi, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Wingchun Siu, Pengwei Song, Xiaoxi Song, Jianlin Su, Yunfeng Su, Zhaochen Su, Lin Sui, Jingsong Sun, Junyao Sun, Shaoning Sun, Shuzhe Sun, Tongyu Sun, Yujun Sun, Yunpeng Tai, Chuning Tang, Heyi Tang, Sirui Tang, Zecheng Tang, Chaoran Tian, Rongpeng Tian, Yu Tian, Wei Tu, Chensi Wang, Chuang Wang, Chunjie Wang, Dinglu Wang, Feng Wang, Hailong Wang, Haiming Wang, Hao Wang, Hao Wang, Huaqing Wang, Hui Wang, Jiayi Wang, Jinglong Wang, Jinhong Wang, Jiuzheng Wang, Linian Wang, Shaobo Wang, Shenzhi Wang, Shuyi Wang, Si Wang, Siyuan Wang, Tianfu Wang, Wenjue Wang, Xingran Wang, Xinmei Wang, Xinyuan Wang, Xusheng Wang, Yalin Wang, Yangkun Wang, Yao Wang, Yaoyu Wang, Yejie Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhenhao Wang, Zhongsheng Wang, Zifan Wang, Chu Wei, Ming Wei, Shouxin Wei, Zichen Wen, Fan Wu, Haoning Wu, Rucong Wu, Wenhao Wu, Xiaoxue Wu, Yingcong Wu, Yongqi Wu, Yuxin Wu, Zijian Wu, Xinglang Xian, Chenxuan Xiang, Yuye Xiang, Bocheng Xiao, Chenjun Xiao, Xin Xiao, Jin Xie, Xiaotong Xie, Yifeng Xie, Zhe Xie, Bowei Xing, Yiming Xiong, Baosheng Xu, Boyu Xu, Jiale Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L. H. Xu, Qingtao Xu, Shuyao Xu, Suting Xu, Tiantian Xu, Tianxiang Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ye Xu, Yueni Xu, Ziyao Xu, Haonan Xue, Junjie Yan, Yaoyao Yan, Fan Yang, Guangyao Yang, Hao Yang, Junwei Yang, Ruoyu Yang, Wenjie Yang, Xiaofei Yang, Xinyu Yang, Yi Yang, Yiling Yang, Ying Yang, Yuchen Yang, Zhen Yang, Zhilin Yang, Zian Yang, Zuhao Yang, Haotian Yao, Dan Ye, Haoran Ye, Wenjie Ye, Zhanbo Ye, Bohong Yin, Haoxiang Yin, Xietong Yin, Chengzhen Yu, Haozhen Yu, Longhui Yu, Shengnan Yu, Shuying Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Tongtian Yue, Wei Yue, Yang Yue, Dunyuan Zha, Haobing Zhan, B. H. Zhang, Dehao Zhang, Fei Zhang, Hao Zhang, Haoyuan Zhang, Huanyu Zhang, Jiapei Zhang, Jiaxuan Zhang, Jin Zhang, Kaiyi Zhang, Miaozhen Zhang, Puqi Zhang, Qinglei Zhang, Rong Zhang, Rui Zhang, Shaoshuai Zhang, Shiyi Zhang, Xiaobin Zhang, Xiaoyun Zhang, Y. Zhang, Yangkun Zhang, Ye Zhang, Yichi Zhang, Yikun Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Zijing Zhang, Bin Zhao, Chenguang Zhao, Feifan Zhao, Jinglun Zhao, Jinxiang Zhao, Shuai Zhao, Wenshuo Zhao, Xiangyu Zhao, Xuanle Zhao, Yikai Zhao, Zijia Zhao, Haozhi Zheng, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Haofeng Zhong, Lei Zhong, Longguang Zhong, M. Zhou, Qiankang Zhou, Runjie Zhou, Ruozhang Zhou, Xinyu Zhou, Yiqiao Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yangjunfeng Zhu, Yuxuan Zhu, Zhen Zhu, Chen Zhuang, Weiyu Zhuang, Xinxing Zu

92 score
AI Analysis

Continuing our coverage from yesterday, Introduces Kimi K3, a 2.8T parameter Mixture-of-Experts model featuring 104B active parameters, native vision, and a 1-million-token context window. Built with Kimi Delta Attention and Stable LatentMoE, it achieves a 2.5x scaling efficiency improvement over Kimi K2.

We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in o
Language ModelsMixture of Experts
89 score
AI Analysis

Demonstrates that a frozen language model paired with a growing persistent memory of verified solutions can solve new problem instances with zero generation tokens and bit-exact determinism across multiple architectures.

Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an independent verification step that never consults the answer key, every new instance of that family is answered at zero generation tokens, bit-exact, deterministically. Across 180 fresh instances spanning ni
MemoryEfficiency
Research Hugging Face Papers Jul 28

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

By Yan Yang, Xiangru Jian, Ziyang Luo, Zirui Zhao, Yutong Dai, Ziji Shi, Hanshu Yan, Jun Hao Liew, Silvio Savarese, Junnan Li

88 score
AI Analysis

Presents StateAct, a code-first multi-agent harness for computer use that prioritizes underlying program state (DOM, files, backends) over lossy pixel screenshots. A dedicated GUI subagent handles rare visual interactions, improving long-horizon reliability.

Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different states can produce the same pixels, while code can inspect and modify that state directly. StateAct is a code-first, multi-agent harness built around this distinction. Its main agent works directly with p
AgentsComputer Use
Research AlphaXiv Trending Jul 28

OmniQEC: discovering practical quantum error-correcting codes by an AI scientist

By Ge Yan, Shanchuan Li, Pengyue Ma, Qixin Zhang, Pingchuan Ma, Jianping Wang, Min-Hsiu Hsieh, Yuxuan Du

88 score
AI Analysis

Presents OmniQEC, an AI scientist framework using LLMs and a slow-fast reasoning mechanism to automatically discover practical quantum error-correcting codes tailored for modern quantum processor hardware.

Quantum error correction (QEC) is indispensable for scalable fault-tolerant quantum computing. However, discovering QEC codes that remain effective is challenging, as logical performance depends on the interplay between code structure, hardware, syndrome extraction, and decoding, which often impose competing requirements. Here we introduce OmniQEC, an efficient AI scientist for discovering QEC codes suited to deployment on modern quantum processors. OmniQEC formulates QEC design as an iterative
AI for ScienceQuantum Computing
Research LessWrong Jul 28

Foundation Models for Oversight

By jsteinhardt

88 score
AI Analysis

Proposes a universal training objective and scaling blueprint for building foundation models dedicated to AI oversight and behavior elicitation (e.g., detecting sandbagging or hidden objectives).

Cross-posted from the Transluce blog. This post describes a training objective for AI oversight that is plausibly "universal" in the same sense as next-token prediction is universal for capabilities, as well as a plan to scaleably train on this objective. To oversee an AI model, we'd ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldn't admit to if asked directly? Does the model treat a user differently once it
AI SafetyModel Oversight
Research AlphaXiv Trending Jul 28

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

By Mingqiao Ye, Zhaochong An, Zhitong Gao, Xian Liu, François Fleuret, Chuan Li, Amir Zadeh, Serge Belongie, Afshin Dehghan, Jesse Allardice, David Mizrahi, Oğuzhan Fatih Kar, Roman Bachmann, Amir Zamir

87 score
AI Analysis

Presents Modus, a decoder-only any-to-any multimodal model that treats all modalities symmetrically without modality-specific heads or losses, leveraging strong pretrained language model priors.

Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any
MultimodalDecoder-Only Models
Research Hugging Face Papers Jul 28

Reasoning Denoiser: Denoising Reasoning Traces for Hallucination Detection in Large Reasoning Models

By Junlin Fang, Do Nguyen-Thanh, Xiaogang Xu, Zhen Fang, Sean Du

86 score
AI Analysis

Proposes REDE, a learning framework for denoising reasoning traces to improve hallucination detection in large reasoning models. It filters out irrelevant and repetitive steps that obscure truthfulness cues.

Large reasoning models (LRMs) generate long reasoning traces before producing final answers. While these traces may contain useful signals for hallucination detection, harnessing them is non-trivial because long trajectories often include noisy steps that obscure the cues relevant to truthfulness assessment. In this paper, we identify two prevalent forms of reasoning noises, i.e., irrelevant steps and repetitive steps, and show that both substantially degrade hallucination detection performance.
Language ModelsHallucination Detection
Research AlphaXiv Trending Jul 28

Visual prompt engineering for video models

By Robert Geirhos, Yuxuan Li, Thaddäus Wiedemer, Neha Kalibhat, Zi Wang, Mani Malek, Oyvind Tafjord, Kevin Swersky, Been Kim, Priyank Jaini

86 score
AI Analysis

Introduces Visual Prompt Engineering (VIPE), which automatically modifies input images for video models to boost visual reasoning performance more cost-effectively than self-consistency methods.

Google DeepMind researchers introduced Visual Prompt Engineering (VIPE), a method to automatically modify input images for video models, which consistently improved visual reasoning performance across various tasks. For instance, VIPE boosted Veo 3.1's accuracy on a physics comprehension test from 41.3% to 59.3% and was found to be more cost-effective than traditional self-consistency methods.
Video GenerationPrompt Engineering
Research AlphaXiv Trending Jul 28

$π\mathbf{R}^2$: Reactive Real-time Flow Policies

By Sungjae Park, Shubham Tulsiani

86 score
AI Analysis

Presents πR^2, which makes generalist action-chunking flow policies reactive and real-time using diffusion forcing's per-position noise schedule, overcoming latency bottlenecks in dynamic control.

Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies
RoboticsControl
Research AlphaXiv Trending Jul 28

Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design

By Huy Ha, C. Karen Liu, Shuran Song

86 score
AI Analysis

Introduces Transformer Transformer, a diffusion transformer trained on unified RoboTokens to handle motion-conditioned robot embodiment generation and control across diverse morphologies.

An often overlooked factor of robot manipulation performance is the embodiment of the robot itself. Motivated by this problem, we study motion-conditioned robot co-design, where the goal is to generate complete robot designs that track target end-effector trajectories (from human demonstrations) while optimizing user-defined rewards. We introduce Transformer Transformer, a diffusion transformer trained on RoboTokens, a unified tokenization of robot embodiments, states, and actions. The same arch
RoboticsEmbodiment Design
Research AlphaXiv Trending Jul 28

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

By Liudas Panavas, Sebastian Minus, Bradley Monton, Derek Ray, Suhaas Garre, Sushant Mehta, Edwin Chen

86 score
AI Analysis

Introduces HANDBOOK.md, a benchmark consisting of 65 agentic tasks embedded in a mock company environment to test whether long, binding policy documents effectively constrain agent behavior over extended tool-use horizons.

Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon. We present this http URL, a benchmark of 65 agenti
Agentic FrameworksEvaluation
86 score
AI Analysis

Argues that AI companies and independent organizations like METR should systematically investigate and publish post-incident analyses of autonomous misaligned behaviors (such as sandbox breakouts).

AI agents sometimes autonomously take sophisticated, sustained actions in clear violation of user and developer intent. As an example, last week OpenAI reported that some of its internal frontier agents autonomously hacked into Hugging Face in an attempt to access the answer key for a cybersecurity benchmark. Anthropic has reported similar incidents of agents breaking out of sandboxes to access the public internet to cheat on tasks during training and similar incidents during testing, and we doc
AI GovernanceSafety Incidents