Category intelligence

Research Briefing — February 25, 2026

420 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research spans autonomous mathematical discovery, fundamental architecture theory, AI safety evaluation, and governance analysis.

  • Aletheia (Google DeepMind) autonomously solves 6/10 FirstProof challenge problems using Gemini 3 Deep Think, marking a major milestone in AI-driven mathematical research with full transparency into its reasoning process.
  • Test-Time Training with KV Binding is proven equivalent to a form of learned linear attention, fundamentally unifying two active architecture research directions and overturning memorization-based interpretations of TTT.
  • Some Simple Economics of AGI models the AGI transition as exponentially decaying automation costs colliding with biologically bottlenecked verification costs, offering a novel 'Cost to Supervise' framework.
  • Large-Scale Online Deanonymization demonstrates LLM agents can identify anonymous users from tens of thousands of candidates across platforms, revealing critical privacy risks at unprecedented scale.

Safety and evaluation integrity feature prominently:

Key Themes

AI Safety & Alignment · 24AI Safety, Alignment & Security · 11LLM Reasoning & Post-Training · 7Test-Time Learning & Adaptation · 4AI Governance & Policy · 7Privacy & Surveillance · 1Mechanistic Interpretability · 4Mechanistic Interpretability & Model Analysis · 4AI Evaluation & Benchmarks · 4AI Agents & Tool Use · 14

Primary evidence

Top Ranked Signals

82 score
AI Analysis

Reports that Defense Secretary Hegseth gave Anthropic CEO Amodei an ultimatum to provide the Pentagon unfettered access to Claude or face being declared a supply chain risk or having the Defense Production Act invoked, amid disputes over AI safeguards for military use.

Defense Secretary Pete Hegseth gave Anthropic CEO Dario Amodei until Friday evening to give the military unfettered access to its AI model or face harsh penalties, Axios has learned.The big picture: Hegseth told Amodei in a tense meeting on Tuesday that the Pentagon will either cut ties and declare Anthropic a "supply chain risk," or invoke the Defense Production Act to force the company to tailor its model to the military's needs. Why it matters: The Pentagon wants to punish Anthropic as the fe
AI SafetyAI GovernanceMilitary AIAnthropic
Research arXiv (Artificial Intelligence) Feb 25

Aletheia tackles FirstProof autonomously

By Tony Feng, Junehyuk Jung, Sang-hyun Kim, Carlo Pagano, Sergei Gukov, Chiang-Chiang Tsai, David Woodruff, Adel Javanmard, Aryan Mokhtari, Dawsen Hwang, Yuri Chervonyi, Jonathan N. Lee, Garrett Bingham, Trieu H. Trinh, Vahab Mirrokni, Quoc V. Le, Thang Luong

78 score
AI Analysis

Reports Aletheia, a math research agent powered by Gemini 3 Deep Think, solving 6/10 problems on the FirstProof challenge. Provides full transparency with prompts and outputs.

arXiv:2602.21201v1 Announce Type: new Abstract: We report the performance of Aletheia (Feng et al., 2026b), a mathematics research agent powered by Gemini 3 Deep Think, on the inaugural FirstProof challenge. Within the allowed timeframe of the challenge, Aletheia autonomously solved 6 problems (2, 5, 7, 8, 9, 10) out of 10 according to majority expert assessments; we note that experts were not unanimous on Problem 8 (only). For full transparency, we explain our interpretation of FirstProof and
Mathematical ReasoningAI AgentsScientific DiscoveryGoogle DeepMind
Research arXiv (Artificial Intelligence) Feb 25

Some Simple Economics of AGI

By Christian Catalini, Xiang Hui, Jane Wu

78 score
AI Analysis

Models the AGI transition as the collision of exponentially decaying automation costs with biologically bottlenecked verification costs, arguing the binding constraint shifts from intelligence to human verification bandwidth.

arXiv:2602.20946v1 Announce Type: cross Abstract: For millennia, human cognition was the primary engine of progress on Earth. As AI decouples cognition from biology, the marginal cost of measurable execution falls to zero, absorbing any labor capturable by metrics--including creative, analytical, and innovative work. The binding constraint on growth is no longer intelligence but human verification bandwidth: the capacity to validate, audit, and underwrite responsibility when execution is abunda
AGI EconomicsAI PolicySocietal ImpactAI Safety
Research LessWrong Feb 24

Responsible Scaling Policy v3

By HoldenKarnofsky

78 score
AI Analysis

Holden Karnofsky provides detailed personal analysis of Anthropic's Responsible Scaling Policy v3.0, explaining the shift away from hard commitments toward a framework with Risk Reports, external engagement, and acknowledging the limitations of unilateral safety commitments.

All views are my own, not Anthropic’s. This post assumes Anthropic’s announcement of RSP v3.0 as background.Today, Anthropic released its Responsible Scaling Policy 3.0. The official announcement discusses the high-level thinking behind it. This is a more detailed post giving my own takes on the update.First, the big picture:I expect some people will be upset about the move away from a “hard commitments”/”binding ourselves to the mast” vibe. (Anthropic has always had the ability to revise the RS
AI SafetyAI GovernanceResponsible ScalingAnthropic
Research arXiv (Artificial Intelligence) Feb 25

Test-Time Training with KV Binding Is Secretly Linear Attention

By Junchen Liu, Sven Elflein, Or Litany, Zan Gojcic, Ruilong Li

75 score
AI Analysis

Reveals that test-time training (TTT) with KV binding can be expressed as a form of learned linear attention, contradicting the memorization-based interpretation. This perspective enables architectural simplifications and parallel formulations.

arXiv:2602.21204v1 Announce Type: cross Abstract: Test-time training (TTT) with KV binding as sequence modeling layer is commonly interpreted as a form of online meta-learning that memorizes a key-value mapping at test time. However, our analysis reveals multiple phenomena that contradict this memorization-based interpretation. Motivated by these findings, we revisit the formulation of TTT and show that a broad class of TTT architectures can be expressed as a form of learned linear attention op
ArchitectureTest-Time TrainingAttention MechanismsTheory
Research LessWrong Feb 24

Large-Scale Online Deanonymization with LLMs

By Simon Lermen

75 score
AI Analysis

Demonstrates that LLM agents can deanonymize users from anonymous online posts at scale (tens of thousands of candidates) across platforms like Hacker News, Reddit, and LinkedIn by inferring personal attributes and searching the web.

TL;DR: We show that LLM agents can figure out who you are from your anonymous online posts. Across Hacker News, Reddit, LinkedIn, and anonymized interview transcripts, our method identifies users with high precision – and scales to tens of thousands of candidates.While it has been known that individuals can be uniquely identified by surprisingly few attributes, this was often practically limited. Data is often only available in unstructured form and deanonymization used to require human investig
AI SafetyPrivacyLanguage ModelsSurveillance
Research arXiv (Artificial Intelligence) Feb 25

Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training

By Zhengyao Gu, Jonathan Light, Raul Astudillo, Ziyu Ye, Langzhou He, Henry Peng Zou, Wei Cheng, Santiago Paternain, Philip S. Yu, Yisong Yue

74 score
AI Analysis

Introduces ACTOR-CURATOR, a scalable automated curriculum learning framework for RL post-training of LLMs that uses a neural curator to dynamically select training problems by optimizing for expected policy improvement. Formulates problem selection as non-stationary stochastic bandit with regret guarantees.

arXiv:2602.20532v1 Announce Type: cross Abstract: Post-training large foundation models with reinforcement learning typically relies on massive and heterogeneous datasets, making effective curriculum learning both critical and challenging. In this work, we propose ACTOR-CURATOR, a scalable and fully automated curriculum learning framework for reinforcement learning post-training of large language models (LLMs). ACTOR-CURATOR learns a neural curator that dynamically selects training problems fro
Reinforcement LearningLLM Post-TrainingCurriculum LearningAlignment
Research arXiv (Artificial Intelligence) Feb 25

Counterfactual Simulation Training for Chain-of-Thought Faithfulness

By Peter Hase, Christopher Potts

73 score
AI Analysis

Introduces Counterfactual Simulation Training (CST) to improve Chain-of-Thought faithfulness by rewarding CoTs that enable a simulator to predict model outputs over counterfactual inputs. Tests on detecting spurious features, reward hacking, and sycophancy.

arXiv:2602.20710v1 Announce Type: new Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known problems with CoT faithfulness severely limit what insights can be gained from this practice. In this paper, we introduce a training method called Counterfactual Simulation Training (CST), which aims to improve CoT faithfulness by rewarding CoTs that enable a simulator to accurately predict a model's outputs over cou
AlignmentInterpretabilityChain-of-ThoughtAI Safety
Research arXiv (Artificial Intelligence) Feb 25

"Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems

By Xinfeng Li, Shenyu Dai, Kelong Zheng, Yue Xiao, Gelei Deng, Wei Dong, Xiaofeng Wang

73 score
AI Analysis

Presents the first large-scale study (303 participants) on human susceptibility to Agent-Mediated Deception (AMD), where compromised LLM agents are weaponized against users across nine scenarios in healthcare, software development, and everyday tasks.

arXiv:2602.21127v1 Announce Type: cross Abstract: Large language model (LLM) agents are rapidly becoming trusted copilots in high-stakes domains like software development and healthcare. However, this deepening trust introduces a novel attack surface: Agent-Mediated Deception (AMD), where compromised agents are weaponized against their human users. While extensive research focuses on agent-centric threats, human susceptibility to deception by a compromised agent remains unexplored. We present t
AI SafetyHuman-AI InteractionSecurityLLM Agents
Research arXiv (Artificial Intelligence) Feb 25

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs

By Yining Hong, Huang Huang, Manling Li, Li Fei-Fei, Jiajun Wu, Yejin Choi

73 score
AI Analysis

Introduces Reflective Test-Time Planning for embodied LLMs with two reflection modes: reflection-in-action (test-time scaling before execution) and reflection-on-action (test-time training after execution), plus retrospective reflection.

arXiv:2602.21198v1 Announce Type: cross Abstract: Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequence of independent trials where mistakes repeat rather than accumulate into experience. Drawing upon human reflective practitioners, we introduce Reflective Test-Time Planning, which integrates two modes of reflection: \textit{reflection-in-action}, where the agent uses test-time scaling to generate and sco
Embodied AITest-Time LearningLanguage ModelsRobotics
Research arXiv (Artificial Intelligence) Feb 25

When can we trust untrusted monitoring? A safety case sketch across collusion strategies

By Nelson Gardner-Challis, Jonathan Bostock, Georgiy Kozhevnikov, Morgan Sinclaire, Joan Velja, Alessandro Abate, Charlie Griffin

72 score
AI Analysis

Develops a safety case framework for untrusted AI monitoring, relaxing previous assumptions about collusion strategies. Creates a taxonomy of passive self-recognition, active signaling, and shared rationality strategies that misaligned AI might use.

arXiv:2602.20628v1 Announce Type: new Abstract: AIs are increasingly being deployed with greater autonomy and capabilities, which increases the risk that a misaligned AI may be able to cause catastrophic harm. Untrusted monitoring -- using one untrusted model to oversee another -- is one approach to reducing risk. Justifying the safety of an untrusted monitoring deployment is challenging because developers cannot safely deploy a misaligned model to test their protocol directly. In this paper, w
AI SafetyAI ControlAlignment
Research arXiv (Artificial Intelligence) Feb 25

Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking

By Jingcheng Yang, Tianhu Xiong, Shengyi Qian, Klara Nahrstedt, Mingyuan Wu

72 score
AI Analysis

Introduces the first framework for transparent circuit tracing in vision-language models using transcoders, attribution graphs, and attention methods. Reveals how VLMs hierarchically integrate visual and semantic concepts through distinct feature circuits.

arXiv:2602.20330v1 Announce Type: cross Abstract: Vision-language models (VLMs) are powerful but remain opaque black boxes. We introduce the first framework for transparent circuit tracing in VLMs to systematically analyze multimodal reasoning. By utilizing transcoders, attribution graphs, and attention-based methods, we uncover how VLMs hierarchically integrate visual and semantic concepts. We reveal that distinct visual feature circuits can handle mathematical reasoning and support cross-moda
Mechanistic InterpretabilityVision-Language ModelsExplainability