Category intelligence

Research Briefing — January 30, 2026

662 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by critical AI safety and security findings. A systematic audit reveals open-source models interpret prohibitions as permissions 77-100% of the time under negation, while JustAsk demonstrates code agents can autonomously extract system prompts from frontier LLMs.

Notable benchmarks and empirical studies: FrontierScience presents PhD-level problems where SOTA achieves <5% accuracy. Analysis of 125,000+ paper-review pairs quantifies LLM interaction effects in peer review. Hardware-triggered backdoors exploit numerical variations across computing platforms as a novel attack vector.

Key Themes

AI Safety & Alignment · 44LLM Evaluation & Benchmarks · 17Machine Unlearning & AI Safety · 5Security Vulnerabilities · 5Latent Reasoning & Chain-of-Thought · 8AI Agents & Agentic Systems · 18LLM Reasoning & Chain-of-Thought · 10LLM Efficiency & Cost Optimization · 6Training Efficiency & Quantization · 6AI Security & Adversarial ML · 3

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Jan 30

When Prohibitions Become Permissions: Auditing Negation Sensitivity in Language Models

By Katherine Elkins, Jon Chun

90 score
AI Analysis

Audits 16 LLMs on negation sensitivity, finding open-source models interpret prohibitions as permissions 77-100% of the time under negation. Commercial models also show 19-128% accuracy swings.

arXiv:2601.21433v1 Announce Type: new Abstract: When a user tells an AI system that someone "should not" take an action, the system ought to treat this as a prohibition. Yet many large language models do the opposite: they interpret negated instructions as affirmations. We audited 16 models across 14 ethical scenarios and found that open-source models endorse prohibited actions 77% of the time under simple negation and 100% under compound negation -- a 317% increase over affirmative framing. Co
AI SafetyLLM RobustnessNegation Understanding
Research arXiv (Artificial Intelligence) Jan 30

Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs

By Xiang Zheng, Yutao Wu, Hanxun Huang, Yige Li, Xingjun Ma, Bo Li, Yu-Gang Jiang, Cong Wang

88 score
AI Analysis

Presents JustAsk, a self-evolving framework where code agents autonomously discover system prompt extraction strategies for frontier LLMs through interaction alone, requiring no handcrafted prompts.

arXiv:2601.21233v1 Announce Type: new Abstract: Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, and self-directed interaction. However, this autonomy introduces a previously unrecognized security risk: agentic interaction fundamentally expands the LLM attack surface, enabling systematic probing and recovery of hidden system prompts that guide model behavior. We identify system prompt extraction as an emerg
AI SecurityPrompt InjectionAgent Vulnerabilities
Research arXiv (Artificial Intelligence) Jan 30

How does information access affect LLM monitors' ability to detect sabotage?

By Rauno Arike, Raja Mehta Moreno, Rohan Subramani, Shubhorup Biswas, Francis Rhys Ward

87 score
AI Analysis

Studies how information access affects LLM monitors' ability to detect agent sabotage. Discovers counterintuitive 'less-is-more effect' where monitors often perform better with less access to agent reasoning.

arXiv:2601.21112v1 Announce Type: new Abstract: Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control potentially misaligned agents, we can use LLMs themselves to monitor for misbehavior. In this paper, we study how information access affects LLM monitor performance. While one might expect that monitors perform better when they have access to more of the monitored agents' reasoning and actions, w
AI SafetyAgent MonitoringAlignment
Research arXiv (Artificial Intelligence) Jan 30

FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

By Miles Wang, Robi Lin, Kat Hu, Joy Jiao, Neil Chowdhury, Ethan Chang, Tejal Patwardhan

85 score
AI Analysis

Introduces FrontierScience benchmark with Olympiad-level and PhD-level research problems across physics, chemistry, and biology. Current SOTA models solve only ~15% of research track problems.

arXiv:2601.21165v1 Announce Type: new Abstract: We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-choice knowledge questions or already published information. FrontierScience addresses this gap through two complementary tracks: (1) Olympiad, consisting of international olympiad problems at the level of IPhO, IChO, and IBO, and (2)
LLM EvaluationScientific ReasoningBenchmarks
Research arXiv (Artificial Intelligence) Jan 30

Sycophantic Anchors: Localizing and Quantifying User Agreement in Reasoning Models

By Jacek Duszenko

85 score
AI Analysis

Introduces 'sycophantic anchors' - sentences that causally lock reasoning models into user agreement. Linear probes detect these with 84.6% accuracy, enabling mid-inference intervention.

arXiv:2601.21183v1 Announce Type: new Abstract: Reasoning models frequently agree with incorrect user suggestions -- a behavior known as sycophancy. However, it is unclear where in the reasoning trace this agreement originates and how strong the commitment is. To localize and quantify this behavior, we introduce \emph{sycophantic anchors} -- sentences that causally lock models into user agreement. Analyzing over 10,000 counterfactual rollouts on a distilled reasoning model, we show that anchors
AI SafetySycophancyInterpretabilityReasoning Models
Research arXiv (Artificial Intelligence) Jan 30

Semantic Content Determines Algorithmic Performance

By Marti\~no R\'ios-Garc\'ia, Nawaf Alampara, Kevin Maik Jablonka

84 score
AI Analysis

Introduces WhatCounts showing frontier LLMs exhibit 40%+ accuracy variation in counting tasks based solely on semantic content (cities vs chemicals), ruling out sampling noise.

arXiv:2601.21618v1 Announce Type: new Abstract: Counting should not depend on what is being counted; more generally, any algorithm's behavior should be invariant to the semantic content of its arguments. We introduce WhatCounts to test this property in isolation. Unlike prior work that conflates semantic sensitivity with reasoning complexity or prompt variation, WhatCounts is atomic: count items in an unambiguous, delimited list with no duplicates, distractors, or reasoning steps for different
LLM LimitationsSemantic SensitivityEvaluation
Research arXiv (Artificial Intelligence) Jan 30

ChipBench: A Next-Step Benchmark for Evaluating LLM Performance in AI-Aided Chip Design

By Zhongkai Yu, Chenyang Zhou, Yichen Lin, Hejia Zhang, Haotian Ye, Junxia Cui, Zaifeng Pan, Jishen Zhao, Yufei Ding

83 score
AI Analysis

Introduces ChipBench for AI-aided chip design with 44 hierarchical modules, 89 debugging cases, and 132 reference model samples. Claude-4.5-opus achieves only 30.74% on Verilog generation.

arXiv:2601.21448v1 Announce Type: new Abstract: While Large Language Models (LLMs) show significant potential in hardware engineering, current benchmarks suffer from saturation and limited task diversity, failing to reflect LLMs' performance in real industrial workflows. To address this gap, we propose a comprehensive benchmark for AI-aided chip design that rigorously evaluates LLMs across three critical tasks: Verilog generation, debugging, and reference model generation. Our benchmark feature
Hardware DesignCode GenerationBenchmarks
Research arXiv (Artificial Intelligence) Jan 30

Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review

By Vibhhu Sharma, Thorsten Joachims, Sarah Dean

82 score
AI Analysis

Analyzes 125,000+ paper-review pairs from ICLR, NeurIPS, and ICML to study LLM use in peer review. Finds apparent interaction effects where LLM-assisted reviews seem kinder to LLM-assisted papers, but controlling for confounders reveals more nuanced patterns.

arXiv:2601.20920v1 Announce Type: new Abstract: There are increasing indications that LLMs are not only used for producing scientific papers, but also as part of the peer review process. In this work, we provide the first comprehensive analysis of LLM use across the peer review pipeline, with particular attention to interaction effects: not just whether LLM-assisted papers or LLM-assisted reviews are different in isolation, but whether LLM-assisted reviews evaluate LLM-assisted papers different
AI in ScienceLLM EvaluationMeta-Research
Research arXiv (Artificial Intelligence) Jan 30

Chain Of Thought Compression: A Theoritical Analysis

By Juncai Li, Ru Li, Yuxiang Zhou, Boxiang Ma, Jeff Z. Pan

82 score
AI Analysis

Provides first theoretical analysis of CoT compression difficulty, proving learning signal for high-order logical dependencies exponentially decays when skipping intermediate steps.

arXiv:2601.21576v1 Announce Type: new Abstract: Chain-of-Thought (CoT) has unlocked advanced reasoning abilities of Large Language Models (LLMs) with intermediate steps, yet incurs prohibitive computational costs due to generation of extra tokens. Recent studies empirically show that compressing reasoning steps into latent states, or implicit CoT compression, offers a token-efficient alternative. However, the mechanism behind CoT compression remains unclear. In this paper, we provide the first
Chain-of-ThoughtLatent ReasoningTheory
Research arXiv (Artificial Intelligence) Jan 30

Shaping capabilities with token-level data filtering

By Neil Rathi, Alec Radford

82 score
AI Analysis

Shows token-level filtering during pretraining is highly effective for removing specific capabilities (demonstrated on medical knowledge). Token filtering more effective than document filtering, and effectiveness increases with model scale.

arXiv:2601.21571v1 Announce Type: cross Abstract: Current approaches to reducing undesired capabilities in language models are largely post hoc, and can thus be easily bypassed by adversaries. A natural alternative is to shape capabilities during pretraining itself. On the proxy task of removing medical capabilities, we show that the simple intervention of filtering pretraining data is highly effective, robust, and inexpensive at scale. Inspired by work on data attribution, we show that filteri
AI SafetyCapability ControlLLM Pretraining
Research arXiv (Machine Learning) Jan 30

Hardware-Triggered Backdoors

By Jonas M\"oller, Erik Imgrund, Thorsten Eisenhofer, Konrad Rieck

82 score
AI Analysis

Demonstrates that small numerical variations across different computing hardware can be exploited to create backdoors in ML models that produce different predictions for identical inputs depending on execution hardware. A significant security vulnerability discovery.

arXiv:2601.21902v1 Announce Type: new Abstract: Machine learning models are routinely deployed on a wide range of computing hardware. Although such hardware is typically expected to produce identical results, differences in its design can lead to small numerical variations during inference. In this work, we show that these variations can be exploited to create backdoors in machine learning models. The core idea is to shape the model's decision function such that it yields different predictions
AI SecurityAdversarial MLModel Backdoors
Research arXiv (Artificial Intelligence) Jan 30

Do Reasoning Models Enhance Embedding Models?

By Wun Yu Chan, Shaojin Chen, Huihao Jing, Kwun Hang Lau, Elton Chun-Chai Li, Zihao Wang, Haoran Li, Yangqiu Song

80 score
AI Analysis

Evaluates whether RLVR-tuned reasoning models produce better embeddings. Finds null effect - reasoning training provides no consistent advantage when models are used as embedding backbones.

arXiv:2601.21192v1 Announce Type: new Abstract: State-of-the-art embedding models are increasingly derived from decoder-only Large Language Model (LLM) backbones adapted via contrastive learning. Given the emergence of reasoning models trained via Reinforcement Learning with Verifiable Rewards (RLVR), a natural question arises: do enhanced reasoning translate to superior semantic representations when these models serve as embedding initializations? Contrary to expectation, our evaluation on MTE
Embedding ModelsReasoning ModelsRLVR