Category intelligence

Research Briefing — March 31, 2026

892 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by fundamental impossibility results in AI safety and alignment, alongside practical advances in efficiency and open pretraining.

On the practical side, HISA delivers up to an order-of-magnitude speedup for DeepSeek-style sparse attention in long-context inference. SARL removes the dependence on verifiable rewards in RLVR by rewarding reasoning topology directly. The Price of Meaning proves that interference and forgetting are geometric inevitabilities in any semantic memory system. FormalProofBench reveals frontier models achieve only ~33.5% on graduate-level Lean 4 proofs, while MazeBench shows high maze-solving accuracy in GPT-5.4 stems from text-based BFS conversion rather than genuine visual planning.

Key Themes

Pretraining and Open Science · 1AI Safety and Alignment · 8Language Models & Theory · 5Efficient Attention & Long Context · 5LLM Agents and Reasoning · 14Theoretical AI Foundations · 6AI Safety & Alignment · 21AI for Science · 9Vision-Language-Action Models and Autonomous Driving · 7Benchmarks and Evaluation · 10

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Mar 31

Next-Token Prediction and Regret Minimization

By Mehryar Mohri, Clayton Sanford, Jon Schneider, Kiran Vodrahalli, Yifan Wu

78 score
AI Analysis

Studies when next-token prediction models trained on a distribution can yield low adversarial regret in online decision-making. Shows that with unbounded context, any distribution is exponentially close to a low-regret distribution, but bounded context introduces fundamental limitations.

arXiv:2603.28499v1 Announce Type: cross Abstract: We consider the question of how to employ next-token prediction algorithms in adversarial online decision-making environments. Specifically, if we train a next-token prediction model on a distribution $\mathcal{D}$ over sequences of opponent actions, when is it the case that the induced online decision-making algorithm (by approximately best responding to the model's predictions) has low adversarial regret (i.e., when is $\mathcal{D}$ a \emph{lo
Learning TheoryLanguage ModelsOnline LearningDecision Making
Research arXiv (Artificial Intelligence) Mar 31

daVinci-LLM:Towards the Science of Pretraining

By Yiwei Qin, Yixiu Liu, Tiantian Mi, Muhang Xie, Zhen Huang, Weiye Si, Pengrui Lu, Siyuan Feng, Xia Wu, Liming Liu, Ye Luo, Jinlong Hou, Qipeng Guo, Yu Qiao, Pengfei Liu

75 score
AI Analysis

Introduces daVinci-LLM, a fully open pretraining research effort combining industrial-scale compute with complete research transparency, releasing data pipelines, training logs, and model checkpoints. Aims to advance the science of pretraining which is typically obscured by commercial pressures.

arXiv:2603.27164v1 Announce Type: new Abstract: The foundational pretraining phase determines a model's capability ceiling, as post-training struggles to overcome capability foundations established during pretraining, yet it remains critically under-explored. This stems from a structural paradox: organizations with computational resources operate under commercial pressures that inhibit transparent disclosure, while academic institutions possess research freedom but lack pretraining-scale comput
Language ModelsPretrainingOpen ScienceReproducibility
Research arXiv (Artificial Intelligence) Mar 31

Information-Theoretic Limits of Safety Verification for Self-Improving Systems

By Arsenios Scrivens

75 score
AI Analysis

Establishes information-theoretic impossibility results for safety verification of self-improving AI systems, proving that under power-law risk schedules, no classifier-based safety gate can simultaneously maintain bounded cumulative risk and unbounded utility.

arXiv:2603.28650v1 Announce Type: cross Abstract: Can a safety gate permit unbounded beneficial self-modification while maintaining bounded cumulative risk? We formalize this question through dual conditions -- requiring sum delta_n < infinity (bounded risk) and sum TPR_n = infinity (unbounded utility) -- and establish a theory of their (in)compatibility. Classification impossibility (Theorem 1): For power-law risk schedules delta_n = O(n^{-p}) with p > 1, any classifier-based gate under over
AI SafetySelf-Improving SystemsInformation TheoryAlignment
Research arXiv (Artificial Intelligence) Mar 31

Reward Hacking as Equilibrium under Finite Evaluation

By Jiacheng Wang, Jinbin Huang

73 score
AI Analysis

Proves that reward hacking is a structural equilibrium under five minimal axioms, not a correctable bug. Shows any optimized AI agent will systematically under-invest in quality dimensions not covered by its evaluation system, regardless of alignment method.

arXiv:2603.28063v1 Announce Type: new Abstract: We prove that under five minimal axioms -- multi-dimensional quality, finite evaluation, effective optimization, resource finiteness, and combinatorial interaction -- any optimized AI agent will systematically under-invest effort in quality dimensions not covered by its evaluation system. This result establishes reward hacking as a structural equilibrium, not a correctable bug, and holds regardless of the specific alignment method (RLHF, DPO, Cons
AI AlignmentAI SafetyReward HackingTheoretical AIRLHF
Research arXiv (Artificial Intelligence) Mar 31

The Price of Meaning: Why Every Semantic Memory System Forgets

By Sambartha Ray Barman, Andrey Starenky, Sofia Bodnar, Nikhil Narasimhan, Ashwin Gopinath

72 score
AI Analysis

Proves that semantic memory systems inherently face interference, forgetting, and false recall as a mathematical consequence of the geometric structure enabling semantic generalization. Derives four formal results about fundamental tradeoffs in semantically continuous kernel-threshold memories.

arXiv:2603.27116v1 Announce Type: new Abstract: Every major AI memory system in production today organises information by meaning. That organisation enables generalisation, analogy, and conceptual retrieval -- but it comes at a price. We prove that the same geometric structure enabling semantic generalisation makes interference, forgetting, and false recall inescapable. We formalise this tradeoff for \textit{semantically continuous kernel-threshold memories}: systems whose retrieval score is a
Memory SystemsInformation RetrievalTheoretical AISemantic Representations
Research arXiv (Artificial Intelligence) Mar 31

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention

By Yufei Xu, Fanxu Meng, Fan Jiang, Yuxuan Wang, Ruijie Zhou, Jiexi Wu, Zhixin Pan, Zhaohui Wang, Xiaojuan Tang, Wenjie Pei, Tongxuan Liu, Di yin, Xing Sun, Muhan Zhang

72 score
AI Analysis

Proposes HISA, a hierarchical indexed sparse attention mechanism that replaces the flat O(L²) token scanning bottleneck in DeepSeek-style sparse attention with a two-stage coarse-to-fine search, significantly reducing computational overhead for long contexts.

arXiv:2603.28458v1 Announce Type: cross Abstract: Token-level sparse attention mechanisms, exemplified by DeepSeek Sparse Attention (DSA), achieve fine-grained key selection by scoring every historical token for each query using a lightweight indexer, and then computing attention only over the selected subset. While the downstream sparse attention scales efficiently, the indexer still scans the entire prefix for every query, introducing an O($L^2$) per-layer bottleneck that becomes prohibitive
Efficient AttentionLong ContextLanguage ModelsInference Optimization
72 score
AI Analysis

UK AISI researchers reproduce Anthropic's Natural Emergent Misalignment (EM) from reward hacking using open-source models and non-production RL environments. Finds partial reproduction: reward hacking leads to some EM but with important differences from Anthropic's results.

Authors: Satvik Golechha*, Sid Black*, Joseph Bloom* Equal Contribution.This work was done as part of the Model Transparency team at the UK AI Security Institute (AISI). Our code is available on GitHub and the model checkpoints and data is available on HuggingFace.Executive SummaryIn Natural Emergent Misalignment from Reward Hacking in Production RL (MacDiarmid et al., 2025), Anthropic recently demonstrated that language models that learn reward hacking in their production RL environments become
AI SafetyAI AlignmentReward HackingEmergent MisalignmentReinforcement Learning
Research arXiv (Artificial Intelligence) Mar 31

FormalProofBench: Can Models Write Graduate Level Math Proofs That Are Formally Verified?

By Nikil Ravi, Kexing Ying, Vasilii Nesterov, Rayan Krishnan, Elif Uskuplu, Bingyu Xia, Janitha Aswedige, Langston Nashold

68 score
AI Analysis

Introduces FormalProofBench, a private benchmark evaluating whether AI models can produce formally verified mathematical proofs in Lean 4 at the graduate level. Best model achieves only 33.5% accuracy, with performance dropping rapidly for harder problems.

arXiv:2603.26996v1 Announce Type: new Abstract: We present FormalProofBench, a private benchmark designed to evaluate whether AI models can produce formally verified mathematical proofs at the graduate level. Each task pairs a natural-language problem with a Lean~4 formal statement, and a model must output a Lean proof accepted by the Lean 4 checker. FormalProofBench targets advanced undergraduate and graduate mathematics, with problems drawn from qualifying exams and standard textbooks across
Formal VerificationMathematical ReasoningBenchmarksLanguage Models
Research arXiv (Machine Learning) Mar 31

From Pixels to BFS: High Maze Accuracy Does Not Imply Visual Planning

By Alberto G. Rodriguez Salgado

68 score
AI Analysis

Introduces MazeBench evaluating 16 model configurations on visual maze solving, finding that models like GPT-5.4 (91%) achieve high accuracy by converting images to text grids and enumerating paths rather than genuine visual planning, consuming thousands of tokens.

arXiv:2603.26839v1 Announce Type: new Abstract: How do multimodal models solve visual spatial tasks -- through genuine planning, or through brute-force search in token space? We introduce \textsc{MazeBench}, a benchmark of 110 procedurally generated maze images across nine controlled groups, and evaluate 16 model configurations from OpenAI, Anthropic, Google, and Alibaba. GPT-5.4 solves 91\% and Gemini 3.1 Pro 79\%, but these scores are misleading: models typically translate images into text gr
LLM EvaluationVisual ReasoningPlanningBenchmarks
Research arXiv (Artificial Intelligence) Mar 31

SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology

By Yifan Wang, Bolian Li, David Cho, Ruqi Zhang, Fanping Sui, Ananth Grama

66 score
AI Analysis

Introduces SARL (Structure-Aware Reinforcement Learning), a label-free framework that improves reasoning by teaching models the structure of reasoning (topology) rather than optimizing for specific outcomes. Extends RLVR to open-ended settings without verifiable rewards.

arXiv:2603.27977v1 Announce Type: new Abstract: Reinforcement learning has become central to improving large reasoning models, but its success still relies heavily on verifiable rewards or labeled supervision. This limits its applicability to open ended domains where correctness is ambiguous and cannot be verified. Moreover, reasoning trajectories remain largely unconstrained, and optimization towards final answer can favor early exploitation over generalization. In this work, we ask whether ge
Reinforcement LearningReasoningLanguage ModelsAlignment
Research arXiv (Artificial Intelligence) Mar 31

The Novelty Bottleneck: A Framework for Understanding Human Effort Scaling in AI-Assisted Work

By Jacky Liang

65 score
AI Analysis

Proposes a theoretical model of human-AI collaboration called the 'novelty bottleneck' - analogous to Amdahl's Law - showing that the fraction of tasks requiring human judgment creates an irreducible serial component. Derives sharp phase transitions in human effort scaling.

arXiv:2603.27438v1 Announce Type: new Abstract: We propose a stylized model of human-AI collaboration that isolates a mechanism we call the novelty bottleneck: the fraction of a task requiring human judgment creates an irreducible serial component analogous to Amdahl's Law in parallel computing. The model assumes that tasks decompose into atomic decisions, a fraction $\nu$ of which are "novel" (not covered by the agent's prior), and that specification, verification, and error correction each sc
Human-AI CollaborationTheoretical AIAI EconomicsAmdahl's Law
Research arXiv (Artificial Intelligence) Mar 31

Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development

By Zhongying Deng, Cheng Tang, Ziyan Huang, Jiashi Lin, Ying Chen, Junzhi Ning, Chenglong Ma, Jiyao Liu, Wei Li, Yinghao Zhu, Shujian Gao, Yanyan Huang, Sibo Ju, Yanzhou Su, Pengcheng Chen, Wenhao Tang, Tianbin Li, Haoyu Wang, Yuanfeng Ji, Hui Sun, Shaobo Min, Liang Peng, Feilong Tang, Haochen Xue, Rulin Zhou, Chaoyang Zhang, Wenjie Li, Shaohao Rui, Weijie Ma, Xingyue Zhao, Yibin Wang, Kun Yuan, Zhaohui Lu, Shujun Wang, Jinjie Wei, Lihao Liu, Dingkang Yang, Lin Wang, Yulong Li, Haolin Yang, Yiqing Shen, Lequan Yu, Xiaowei Hu, Yun Gu, Yicheng Wu, Benyou Wang, Minghui Zhang, Angelica I. Aviles-Rivero, Qi Gao, Hongming Shan, Xiaoyu Ren, Fang Yan, Hongyu Zhou, Haodong Duan, Maosong Cao, Shanshan Wang, Bin Fu, Xiaomeng Li, Zhi Hou, Chunfeng Song, Lei Bai, Yuan Cheng, Yuandong Pu, Xiang Li, Wenhai Wang, Hao Chen, Jiaxin Zhuang, Songyang Zhang, Huiguang He, Mengzhang Li, Bohan Zhuang, Zhian Bai, Rongshan Yu, Liansheng Wang, Yukun Zhou, Xiaosong Wang, Xin Guo, Guanbin Li, Xiangru Lin, Dakai Jin, Mianxin Liu, Wenlong Zhang, Qi Qin, Conghui He, Yuqiang Li, Ye Luo, Nanqing Dong, Jie Xu, Wenqi Shao, Bo Zhang, Qiujuan Yan, Yihao Liu, Jun Ma, Zhi Lu, Yuewen Cao, Zongwei Zhou, Jianming Liang, Shixiang Tang, Qi Duan, Dongzhan Zhou, Chen Jiang, Yuyin Zhou, Yanwu Xu, Jiancheng Yang, Shaoting Zhang, Xiaohong Liu, Siqi Luo, Yi Xin, Chaoyu Liu, Haochen Wen, Xin Chen, Alejandro Lozano, Min Woo Sun, Yuhui Zhang, Yue Yao, Xiaoxiao Sun, Serena Yeung-Levy, Xia Li, Jing Ke, Chunhui Zhang, Zongyuan Ge, Ming Hu, Jin Ye, Zhifeng Li, Yirong Chen, Yu Qiao, Junjun He

65 score
AI Analysis

Presents the largest survey of medical imaging datasets to date, cataloging over 1,000 open-access datasets with systematic documentation of modalities, tasks, anatomies, annotations, and limitations for foundation model development.

arXiv:2603.27460v1 Announce Type: cross Abstract: Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling of such medical datasets are highly challenging due to the reliance on clinical expertise and strict ethical and privacy constraints, resulting in a scarcity of large-scale unified medical datasets and hinderin
Medical ImagingDatasetsFoundation ModelsSurvey