Category intelligence

Research Briefing — January 21, 2026

1056 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research reveals critical challenges to conventional training wisdom and exposes multiple safety vulnerabilities across deployed systems.

Training Paradigm Reassessment:

Safety & Security Vulnerabilities:

Reasoning Model Insights:

Architecture Innovation: Threshold Differential Attention eliminates attention sinks while achieving ultra-sparsity and improved long-context robustness.

Key Themes

LLM Training & Fine-tuning · 9AI Safety & Alignment · 40Foundation Models · 2LLM Safety and Security · 14LLM Reasoning & Evaluation · 14AI Safety & Security · 20AI Safety & Robustness · 7Efficient Architectures · 6LLM Reasoning and Chain-of-Thought · 12MLLM/VLM Evaluation & Reasoning · 9

Primary evidence

Top Ranked Signals

Research arXiv (Machine Learning) Jan 21

Do Instruction-Tuned Models Always Perform Better Than Base Models? Evidence from Math and Domain-Shifted Benchmarks

By Prateek Munjal, Clement Christophe, Ronnie Rajan, Praveenkumar Kanithi

88 score
AI Analysis

Investigates whether instruction-tuned models always outperform base models, finding that base models consistently outperform instruction-tuned variants in zero-shot CoT settings on GSM8K (drops up to 32.67% for Llama3-70B). Instruction tuning appears to induce pattern matching rather than genuine reasoning improvement.

arXiv:2601.13244v1 Announce Type: new Abstract: Instruction finetuning is standard practice for improving LLM performance, yet it remains unclear whether it enhances reasoning or merely induces surface-level pattern matching. We investigate this by evaluating base and instruction-tuned models on standard math benchmarks, structurally perturbed variants, and domain-shifted tasks. Our analysis highlights two key (often overlooked) limitations of instruction tuning. First, the performance advantag
LLM TrainingInstruction TuningReasoning
Research arXiv (Machine Learning) Jan 21

Balancing Classification and Calibration Performance in Decision-Making LLMs via Calibration Aware Reinforcement Learning

By Duygu Nur Yaldiz, Evangelia Spiliopoulou, Zheng Qi, Siddharth Varia, Srikanth Doss, Nikolaos Pappas

85 score
AI Analysis

Systematic study showing RLVR improves task performance but produces extremely overconfident models, while SFT yields better calibration even under distribution shift. Proposes calibration-aware RL approach to balance classification and calibration.

arXiv:2601.13284v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in decision-making tasks, where not only accuracy but also reliable confidence estimates are essential. Well-calibrated confidence enables downstream systems to decide when to trust a model and when to defer to fallback mechanisms. In this work, we conduct a systematic study of calibration in two widely used fine-tuning paradigms: supervised fine-tuning (SFT) and reinforcement learning with ve
AI SafetyCalibrationReinforcement LearningLLM Training
Research arXiv (Computer Vision) Jan 21

CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning

By Wenxin Ma, Chenlong Wang, Ruisheng Yuan, Hao Chen, Nanru Dai, S. Kevin Zhou, Yijun Yang, Alan Yuille, Jieneng Chen

85 score
AI Analysis

Introduces CausalSpatial benchmark evaluating whether MLLMs can anticipate consequences of object motions across collision, compatibility, occlusion and trajectory tasks. Reveals severe gap: humans score 84% while GPT-5 achieves only 54%, exposing over-reliance on textual reasoning.

arXiv:2601.13304v1 Announce Type: new Abstract: Humans can look at a static scene and instantly predict what happens next -- will moving this object cause a collision? We call this ability Causal Spatial Reasoning. However, current multimodal large language models (MLLMs) cannot do this, as they remain largely restricted to static spatial perception, struggling to answer "what-if" questions in a 3D scene. We introduce CausalSpatial, a diagnostic benchmark evaluating whether models can anticipat
BenchmarksMultimodal ReasoningCausal ReasoningMLLM Evaluation
Research arXiv (Artificial Intelligence) Jan 21

Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents?

By Yi Qian, Kunwei Qian, Xingbang He, Ligeng Chen, Jikang Zhang, Tiantai Zhang, Haiyang Wei, Linzhang Wang, Hao Wu, Bing Mao

82 score
AI Analysis

Discovers 'Action Rebinding' - a critical security vulnerability in multimodal GUI agents where zero-permission apps can hijack agent actions by exploiting the gap between observation and action execution. Demonstrates that Visual Atomicity assumption is invalid on Android.

arXiv:2601.12349v1 Announce Type: cross Abstract: Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted with perceiving screen content and injecting inputs. However, their design operates under the implicit assumption of Visual Atomicity: that the UI state remains invariant between observation and action. We demonstrate that this assumption is fundamentally invalid in Android, creating a critical attack surface. We present Action Reb
AI SafetySecurity VulnerabilitiesAgentic AIMultimodal Models
Research arXiv (Artificial Intelligence) Jan 21

Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs

By Jackson Kaunismaa, Avery Griffin, John Hughes, Christina Q. Knight, Mrinank Sharma, Erik Jones

82 score
AI Analysis

Demonstrates that safeguarded frontier models can be used to elicit harmful capabilities in open-source models through three-stage elicitation attacks using adjacent-domain prompts that bypass safeguards.

arXiv:2601.13528v1 Announce Type: cross Abstract: Model developers implement safeguards in frontier models to prevent misuse, for example, by employing classifiers to filter dangerous outputs. In this work, we demonstrate that even robustly safeguarded models can be used to elicit harmful capabilities in open-source models through elicitation attacks. Our elicitation attacks consist of three stages: (i) constructing prompts in adjacent domains to a target harmful task that do not request danger
AI SafetyModel SecurityCapability Elicitation
Research arXiv (Artificial Intelligence) Jan 21

APEX-Agents

By Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, Julien Benchek, David Ostrofsky, Anirudh Ravichandran, Debnil Sur, Neel Venugopal, Alannah Hsia, Isaac Robinson, Calix Huang, Olivia Varones, Daniyal Khan, Michael Haines, Zach Richards, Chirag Mahapatra, Brendan Foody, Osvald Nitski

82 score
AI Analysis

APEX-Agents benchmarks AI agents on professional tasks from investment banking, consulting, and law with 480 long-horizon cross-application tasks. Gemini 3 Flash achieves 24%, followed by GPT-5.2 and Claude Opus 4.5.

arXiv:2601.14242v1 Announce Type: cross Abstract: We introduce the AI Productivity Index for Agents (APEX-Agents), a benchmark for assessing whether AI agents can execute long-horizon, cross-application tasks created by investment banking analysts, management consultants, and corporate lawyers. APEX-Agents requires agents to navigate realistic work environments with files and tools. We test eight agents for the leaderboard using Pass@1. Gemini 3 Flash (Thinking=High) achieves the highest score
AI AgentsBenchmarksProfessional AIEvaluation
Research arXiv (Machine Learning) Jan 21

A Comprehensive Evaluation of LLM Reasoning: From Single-Model to Multi-Agent Paradigms

By Yapeng Li, Jiakuo Yu, Zhixin Liu, Xinnan Liu, Jing Yu, Songze Li, Tonghua Su

82 score
AI Analysis

Comprehensive unified evaluation of LLM reasoning paradigms from single-model to multi-agent systems, characterizing performance across benchmarks and analyzing cost-accuracy trade-offs. Probes role-specific capability demands in MAS.

arXiv:2601.13243v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as reasoning systems, where reasoning paradigms - such as Chain-of-Thought (CoT) and multi-agent systems (MAS) - play a critical role, yet their relative effectiveness and cost-accuracy trade-offs remain poorly understood. In this work, we conduct a comprehensive and unified evaluation of reasoning paradigms, spanning direct single-model generation, CoT-augmented single-model reasoning, and re
LLM ReasoningMulti-Agent SystemsBenchmarking
Research arXiv (Computer Vision) Jan 21

Human detectors are surprisingly powerful reward models

By Kumar Ashutosh, XuDong Wang, Xi Yin, Kristen Grauman, Adam Polyak, Ishan Misra, Rohit Girdhar

82 score
AI Analysis

Proposes HuDA, a surprisingly simple reward model using human detection confidence and temporal prompt alignment to improve human motion in generated videos. Off-the-shelf models outperform specialized methods without training.

arXiv:2601.14037v1 Announce Type: new Abstract: Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when synthesizing humans performing dynamic actions such as sports, dance, etc. Generated videos often exhibit missing or extra limbs, distorted poses, or physically implausible actions. In this work, we propose a remarkably simple reward model, HuDA, to quantify and improve th
Video GenerationReward ModelsHuman Motion SynthesisRLHF
Research arXiv (Artificial Intelligence) Jan 21

AI-generated data contamination erodes pathological variability and diagnostic reliability

By Hongyu He, Shaowen Xiang, Ye Zhang, Yingtao Zhu, Jin Zhang, Hao Deng, Emily Alsentzer, Qingyu Chen, Kun-Hsing Yu, Andrew Marmenshall, Tingting Chen, Srinivas Anumasa, Daniel Ebner, Dean Ho, Kee Yuan Ngiam, Ching-Yu Cheng, Dianbo Liu

80 score
AI Analysis

Demonstrates that AI-generated data contamination in medical AI creates feedback loop causing erosion of pathological variability and diagnostic reliability, with models converging toward generic phenotypes regardless of architecture.

arXiv:2601.12946v1 Announce Type: cross Abstract: Generative artificial intelligence (AI) is rapidly populating medical records with synthetic content, creating a feedback loop where future models are increasingly at risk of training on uncurated AI-generated data. However, the clinical consequences of this AI-generated data contamination remain unexplored. Here, we show that in the absence of mandatory human verification, this self-referential cycle drives a rapid erosion of pathological varia
Medical AIAI SafetyData ContaminationSynthetic Data
Research arXiv (Machine Learning) Jan 21

Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling

By Xingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu, Neil Shah, Tong Zhao

80 score
AI Analysis

Proposes Threshold Differential Attention (TDA), a sink-free attention mechanism achieving ultra-sparsity and improved robustness at longer sequence lengths. Uses row-wise extreme-value thresholding with length-dependent gating without computational overhead of projection methods.

arXiv:2601.12145v1 Announce Type: new Abstract: Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, and probability mass disperses as sequence lengths increase. We tackle these problems with Threshold Differential Attention (TDA), a sink-free attention mechanism that achieves ultra-sparsity and improved robustness at longer sequence lengths without the computational overhead of projection met
Attention MechanismsLong ContextEfficient Architectures
Research arXiv (Machine Learning) Jan 21

Dissecting Linear Recurrent Models: How Different Gating Strategies Drive Selectivity and Generalization

By Younes Bouhadjar, Maxime Fabre, Felix Schmidt, Emre Neftci

79 score
AI Analysis

Proposes refined taxonomy of linear recurrent models (Mamba-like architectures) and introduces SelectivBench, lightweight benchmarks revealing how different gating strategies drive selectivity and generalization. Enables systematic comparison without resource-intensive experiments.

arXiv:2601.12598v1 Announce Type: new Abstract: Linear recurrent neural networks have emerged as efficient alternatives to the original Transformer's softmax attention mechanism, thanks to their highly parallelizable training and constant memory and computation requirements at inference. Iterative refinements of these models have introduced an increasing number of architectural mechanisms, leading to increased complexity and computational costs. Nevertheless, systematic direct comparisons among
Efficient ArchitecturesState Space ModelsBenchmarking
Research arXiv (Computation and Language) Jan 21

The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning

By Renmiao Chen, Yida Lu, Shiyao Cui, Xuan Ouyang, Victor Shea-Jay Huang, Shumin Zhang, Chengwei Pan, Han Qiu, Minlie Huang

79 score
AI Analysis

MIR-SafetyBench is the first benchmark focused on multi-image reasoning safety in MLLMs (2,676 instances, 9 image relations). Reveals troubling trend: models with better multi-image reasoning are MORE vulnerable to safety exploits.

arXiv:2601.14127v1 Announce Type: cross Abstract: As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi-image instructions, this advancement may pose new safety risks. We study this problem by introducing MIR-SafetyBench, the first benchmark focused on multi-image reasoning safety, which consists of 2,676 instances across a taxonomy of 9 multi-image relations. Our extensive evaluations on 19 MLLMs reveal a troubling trend: models with more
AI SafetyMultimodal LLMsBenchmarksAdversarial Robustness