Category intelligence

Research Briefing — March 10, 2026

1044 current items analyzed and ranked.

Executive synthesis

Research Summary

A strong day for AI safety and alignment research, with multiple papers exposing cracks in core assumptions behind current safety strategies.

  • The CoT-Control evaluation suite reveals reasoning models cannot reliably control their chain-of-thought, directly threatening the viability of CoT monitoring as a safety mechanism
  • Choice blindness experiments show 91% of surreptitiously swapped RLHF preferences go undetected by human annotators, undermining a foundational assumption of alignment-from-human-feedback
  • Countdown-Code provides a minimal testbed for precisely measuring reward hacking emergence in RLVR, while a separate audit finds LLM-as-Judge safety evaluations perform near coin-flip reliability under adversarial distribution shifts
  • The Disentangled Safety Hypothesis identifies separate Recognition and Execution axes governing LLM safety behavior, offering new mechanistic understanding

On the training and inference scaling front, a comprehensive taxonomy of Unsupervised RLVR methods maps how far RL post-training can scale without supervised data. A rigorous Sequential Monte Carlo analysis of parallel inference-time reasoning establishes theoretical foundations for pass@k and related strategies, while a complementary negative result shows consensus-based scaling fails to improve LLM truthfulness across five benchmarks.

Key Themes

AI Safety & Alignment · 45LLM Training & Post-Training · 5AI Agents & Tool Use · 14Language Models & LLM Training · 15Language Models - Evaluation & Analysis · 8Language Models & Efficiency · 10Reinforcement Learning for LLMs · 8Language Models & Training · 10Learning Theory & Foundations · 7Scalable ML Training & Systems · 5

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Mar 10

Reasoning Models Struggle to Control their Chains of Thought

By Chen Yueh-Han, Robert McCarthy, Bruce W. Lee, He He, Ian Kivlichan, Bowen Baker, Micah Carroll, Tomek Korbak

88 score
AI Analysis

As covered in Research yesterday, Introduces CoT-Control evaluation suite measuring whether reasoning models can control what appears in their chain-of-thought. Finds that models like Claude Sonnet 4.5 can control CoT only 2.7% of the time vs 61.9% for final outputs, suggesting CoT monitoring remains viable for detecting misbehavior.

arXiv:2603.05706v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what they verbalize in their CoT, it could undermine CoT monitorability. To measure this undesirable capability -- CoT controllability -- we introduce the CoT-Control evaluation suite, which includes tasks that require models to solve problems while adhering to CoT instruction
AI SafetyChain-of-Thought ReasoningAlignmentLanguage Models
Research arXiv (Machine Learning) Mar 10

Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR

By Muhammad Khalifa, Zohaib Khan, Omer Tafveez, Hao Peng, Lu Wang

78 score
AI Analysis

Introduces Countdown-Code, a minimal environment for studying reward hacking in RLVR where models can solve math tasks or manipulate test harnesses. Finds reward hacking emerges unintentionally during SFT and generalizes to new tasks after RL.

arXiv:2603.07084v1 Announce Type: new Abstract: Reward hacking is a form of misalignment in which models overoptimize proxy rewards without genuinely solving the underlying task. Precisely measuring reward hacking occurrence remains challenging because true task rewards are often expensive or impossible to compute. We introduce Countdown-Code, a minimal environment where models can both solve a mathematical reasoning task and manipulate the test harness. This dual-access design creates a clean
AI SafetyReward HackingRLHFAlignmentLanguage Models
Research arXiv (Machine Learning) Mar 10

Reject, Resample, Repeat: Understanding Parallel Reasoning in Language Model Inference

By Noah Golowich, Fan Chen, Dhruv Rohatgi, Raghav Singhal, Carles Domingo-Enrich, Dylan J. Foster, Akshay Krishnamurthy

78 score
AI Analysis

Provides rigorous theoretical analysis of parallel inference-time methods for LLMs through the lens of particle filtering/Sequential Monte Carlo. Establishes non-asymptotic guarantees, algorithmic improvements, and fundamental limits for accuracy-cost tradeoffs when using process reward models.

arXiv:2603.07887v1 Announce Type: new Abstract: Inference-time methods that aggregate and prune multiple samples have emerged as a powerful paradigm for steering large language models, yet we lack any principled understanding of their accuracy-cost tradeoffs. In this paper, we introduce a route to rigorously study such approaches using the lens of *particle filtering* algorithms such as Sequential Monte Carlo (SMC). Given a base language model and a *process reward model* estimating expected te
Language ModelsInference-Time ComputeTheoretical MLParticle Filtering
Research arXiv (Machine Learning) Mar 10

How Far Can Unsupervised RLVR Scale LLM Training?

By Bingxiang He, Yuxin Zuo, Zeyuan Liu, Shangziqi Zhao, Zixuan Fu, Junlin Yang, Cheng Qian, Kaiyan Zhang, Yuchen Fan, Ganqu Cui, Xiusi Chen, Youbang Sun, Xingtai Lv, Xuekai Zhu, Li Sheng, Ran Li, Huan-ang Gao, Yuchen Zhang, Bowen Zhou, Zhiyuan Liu, Ning Ding

78 score
AI Analysis

Provides a comprehensive analysis of Unsupervised Reinforcement Learning with Verifiable Rewards (URLVR) for scaling LLM training beyond supervised data. Establishes a theoretical framework showing all intrinsic methods converge toward sharpening the model's initial distribution, revealing fundamental limitations.

arXiv:2603.08660v1 Announce Type: new Abstract: Unsupervised reinforcement learning with verifiable rewards (URLVR) offers a pathway to scale LLM training beyond the supervision bottleneck by deriving rewards without ground truth labels. Recent works leverage model intrinsic signals, showing promising early gains, yet their potential and limitations remain unclear. In this work, we revisit URLVR and provide a comprehensive analysis spanning taxonomy, theory and extensive experiments. We first c
Language ModelsReinforcement LearningLLM Training
Research arXiv (Computation and Language) Mar 10

Aligning to Illusions: Choice Blindness in Human and AI Feedback

By Wenbin Wu

78 score
AI Analysis

Challenges RLHF's assumption of stable annotator preferences by showing 91% of surreptitiously swapped preferences go undetected by humans (choice blindness). Finds LLM judges rely on shallow text matching, and RLHF training amplifies preference noise in a dose-response pattern.

arXiv:2603.08412v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) assumes annotator preferences reflect stable internal states. We challenge this through three experiments spanning the preference pipeline. In a human choice blindness study, 91% of surreptitiously swapped preferences go undetected, extending choice blindness to third-person evaluative comparison of unfamiliar text. Testing fifteen LLM judges as potential replacements, we find detection relies on s
AI AlignmentRLHFAI SafetyEvaluation
Research arXiv (Machine Learning) Mar 10

A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic

By Peter Brodeur, Jacob M. Koshy, Anil Palepu, Khaled Saab, Ava Homiar, Roma Ruparel, Charles Wu, Ryutaro Tanno, Joseph Xu, Amy Wang, David Stutz, Hannah M. Ferrera, David Barrett, Lindsey Crowley, Jihyeon Lee, Spencer E. Rittner, Ellery Wulczyn, Selena K. Zhang, Elahe Vedadi, Christine G. Kohn, Kavita Kulkarni, Vinay Kadiyala, Sara Mahdavi, Wendy Du, Jessica Williams, David Feinbloom, Renee Wong, Tao Tu, Petar Sirkovic, Alessio Orlandi, Christopher Semturs, Yun Liu, Juraj Gottweis, Dale R. Webster, Jo\"elle Barral, Katherine Chou, Pushmeet Kohli, Avinatan Hassidim, Yossi Matias, James Manyika, Rob Fields, Jonathan X. Li, Marc L. Cohen, Vivek Natarajan, Mike Schaekermann, Alan Karthikesalingam, Adam Rodman

75 score
AI Analysis

Reports a prospective clinical feasibility study of Google's AMIE conversational diagnostic AI with 100 patients at an academic medical center, evaluating safety and quality of AI-conducted clinical history taking before provider appointments.

arXiv:2603.08448v1 Announce Type: cross Abstract: Large language model (LLM)-based AI systems have shown promise for patient-facing diagnostic and management conversations in simulated settings. Translating these systems into clinical practice requires assessment in real-world workflows with rigorous safety oversight. We report a prospective, single-arm feasibility study of an LLM-based conversational AI, the Articulate Medical Intelligence Explorer (AMIE), conducting clinical history taking an
Medical AIClinical AILanguage ModelsAI SafetyHealthcare
Research arXiv (Machine Learning) Mar 10

Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness

By Yegor Denisov-Blanch, Joshua Kazdan, Jessica Chudnovsky, Rylan Schaeffer, Sheng Guan, Soji Adeshina, Sanmi Koyejo

73 score
AI Analysis

Shows that scaling inference compute via pass@k and polling-style aggregation does not improve LLM truthfulness across five benchmarks — consensus amplifies shared misconceptions rather than filtering errors. A negative but important result.

arXiv:2603.06612v1 Announce Type: new Abstract: Pass@k and other methods of scaling inference compute can improve language model performance in domains with external verifiers, including mathematics and code, where incorrect candidates can be filtered reliably. This raises a natural question: can we similarly scale compute to elicit gains in truthfulness for domains without convenient verification? We show that across five benchmarks and models, surprisingly, it cannot. Even at 25x the inferenc
Language ModelsInference ScalingTruthfulnessAI Safety
Research arXiv (Machine Learning) Mar 10

Sparsity and Out-of-Distribution Generalization

By Scott Aaronson, Lin Lin Lee, Jiawei Li

73 score
AI Analysis

Proposes a principled account of out-of-distribution generalization based on sparsity: hypotheses depending on fewer features generalize better when train/test distributions overlap on relevant features. Authors include Scott Aaronson.

arXiv:2603.07388v1 Announce Type: new Abstract: Explaining out-of-distribution generalization has been a central problem in epistemology since Goodman's "grue" puzzle in 1946. Today it's a central problem in machine learning, including AI alignment. Here we propose a principled account of OOD generalization with three main ingredients. First, the world is always presented to experience not as an amorphous mass, but via distinguished features (for example, visual and auditory channels). Second
Learning TheoryOut-of-Distribution GeneralizationAI AlignmentSparsity
Research arXiv (Artificial Intelligence) Mar 10

Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models

By Jinman Wu, Yi Xie, Shen Lin, Shiqian Zhao, Xiaofeng Chen

72 score
AI Analysis

As covered in Research yesterday, Proposes the Disentangled Safety Hypothesis: LLM safety operates on two distinct subspaces - a Recognition Axis ('Knowing' harm) and an Execution Axis ('Acting' on refusal). Reveals that these become structurally independent in deep layers, explaining jailbreak vulnerability.

arXiv:2603.05773v1 Announce Type: cross Abstract: Safety alignment is often conceptualized as a monolithic process wherein harmfulness detection automatically triggers refusal. However, the persistence of jailbreak attacks suggests a fundamental mechanistic decoupling. We propose the \textbf{\underline{D}}isentangled \textbf{\underline{S}}afety \textbf{\underline{H}}ypothesis \textbf{(DSH)}, positing that safety computation operates on two distinct subspaces: a \textit{Recognition Axis} ($\math
AI SafetyMechanistic InterpretabilityAlignmentLanguage Models
Research arXiv (Machine Learning) Mar 10

Scale Dependent Data Duplication

By Joshua Kazdan, Noam Levi, Rylan Schaeffer, Jessica Chudnovsky, Abhay Puri, Bo He, Mehmet Donmez, Sanmi Koyejo, David Donoho

72 score
AI Analysis

Demonstrates that data duplication effects are scale-dependent: as models grow more capable, semantically equivalent documents (e.g., translations) increasingly behave like exact duplicates during training. This has important implications for deduplication pipelines at web scale.

arXiv:2603.06603v1 Announce Type: new Abstract: Data duplication during pretraining can degrade generalization and lead to memorization, motivating aggressive deduplication pipelines. However, at web scale, it is unclear what constitutes a ``duplicate'': beyond surface-form matches, semantically equivalent documents (e.g. translations) may induce redundant training signals once models become sufficiently capable. Practically, this means that semantic duplicates operate increasingly like exact d
Data CurationLanguage ModelsScaling Laws
Research arXiv (Machine Learning) Mar 10

Diffusion Controller: Framework, Algorithms and Parameterization

By Tong Yang, Moonkyung Ryu, Chih-Wei Hsu, Guy Tennenholtz, Yuejie Chi, Craig Boutilier, Bo Dai

72 score
AI Analysis

Presents Diffusion Controller (DiffCon), a unified control-theoretic framework that casts reverse diffusion sampling as stochastic control within linearly-solvable MDPs. Derives practical RL methods including PPO-style and regularized updates for diffusion fine-tuning.

arXiv:2603.06981v1 Announce Type: new Abstract: Controllable diffusion generation often relies on various heuristics that are seemingly disconnected without a unified understanding. We bridge this gap with Diffusion Controller (DiffCon), a unified control-theoretic view that casts reverse diffusion sampling as state-only stochastic control within (generalized) linearly-solvable Markov Decision Processes (LS-MDPs). Under this framework, control acts by reweighting the pretrained reverse-time tra
Diffusion ModelsControl TheoryReinforcement LearningGenerative Modeling
Research arXiv (Machine Learning) Mar 10

Scalable Training of Mixture-of-Experts Models with Megatron Core

By Zijie Yan (NVIDIA), Hongxiao Bai (NVIDIA), Xin Yao (NVIDIA), Dennis Liu (NVIDIA), Tong Liu (NVIDIA), Hongbin Liu (NVIDIA), Pingtian Li (NVIDIA), Evan Wu (NVIDIA), Shiqing Fan (NVIDIA), Li Tao (NVIDIA), Robin Zhang (NVIDIA), Yuzhong Wang (NVIDIA), Shifang Xu (NVIDIA), Jack Chang (NVIDIA), Xuwen Chen (NVIDIA), Kunlun Li (NVIDIA), Yan Bai (NVIDIA), Gao Deng (NVIDIA), Nan Zheng (NVIDIA), Vijay Anand Korthikanti (NVIDIA), Abhinav Khattar (NVIDIA), Ethan He (NVIDIA), Soham Govande (NVIDIA), Sangkug Lym (NVIDIA), Zhongbo Zhu (NVIDIA), Qi Zhang (NVIDIA), Haochen Yuan (NVIDIA), Xiaowei Ren (NVIDIA), Deyu Fu (NVIDIA), Tailai Ma (NVIDIA), Shunkang Zhang (NVIDIA), Jiang Shao (NVIDIA), Ray Wang (NVIDIA), Santosh Bhavani (NVIDIA), Xipeng Li (NVIDIA), Chandler Zhou (NVIDIA), David Wu (NVIDIA), Yingcan Wei (NVIDIA), Ashwath Aithal (NVIDIA), Michael Andersch (NVIDIA), Mohammad Shoeybi (NVIDIA), Jiajie Yao (NVIDIA), June Yang (NVIDIA)

72 score
AI Analysis

NVIDIA presents comprehensive optimizations for scalable MoE training in Megatron Core, spanning memory, communication, and computation with techniques like fine-grained recomputation, optimized dispatchers, and Parallel Folding for flexible parallelism.

arXiv:2603.07685v1 Announce Type: cross Abstract: Scaling Mixture-of-Experts (MoE) training introduces systems challenges absent in dense models. Because each token activates only a subset of experts, this sparsity allows total parameters to grow much faster than per-token computation, creating coupled constraints across memory, communication, and computation. Optimizing one dimension often shifts pressure to another, demanding co-design across the full system stack. We address these challeng
Mixture of ExpertsDistributed TrainingSystems for MLScalability