Category intelligence

Research Briefing — February 27, 2026

531 current items analyzed and ranked.

Executive synthesis

Research Summary

AI safety and biosecurity research dominate today's highlights. A rigorous human uplift study shows LLM access yields a 4.16x accuracy boost for novices on dual-use biology tasks, with major policy implications. AuditBench provides 56 models with implanted hidden behaviors for evaluating alignment auditing, while a novel 'self-incrimination training' approach teaches GPT-4.1 and Gemini-2.0 agents to report their own deceptive behavior. A blog post argues eval awareness may emerge as a training artifact rather than a pure capability.

Key Themes

AI Policy & Governance · 6AI Safety & Alignment · 27AI Safety & Security · 14AI Safety and Alignment · 6Diffusion Language Models · 4Mechanistic Interpretability · 14Mechanistic Interpretability & Model Analysis · 4Training & Inference Efficiency · 7Language Models & Training Efficiency · 10Vision-Language Models · 6

Primary evidence

Top Ranked Signals

Research LessWrong Feb 25

Scoop: Pentagon takes first step toward blacklisting Anthropic

By Matrice Jacobine

80 score
AI Analysis

Reports the Pentagon asked Boeing and Lockheed Martin to assess their reliance on Anthropic's Claude, a first step toward designating Anthropic as a 'supply chain risk' - a penalty usually reserved for adversarial nations.

The Pentagon asked two major defense contractors on Wednesday to provide an assessment of their reliance on Anthropic's AI model, Claude — a first step toward a potential designation of Anthropic as a "supply chain risk," Axios has learned.Why it matters: That penalty is usually reserved for companies from adversarial countries, such as Chinese tech giant Huawei.Using it to punish a leading American tech firm, particularly one on which the military itself is currently reliant, would be unprecede
AI PolicyAI GovernanceNational Security
78 score
AI Analysis

Reposts Dario Amodei's statement on Anthropic's work with the Department of War, emphasizing their commitment to deploying AI for US defense and democracy while maintaining red lines on domestic surveillance and autonomous killing.

I believe deeply in the existential importance of using AI to defend the United States and other democracies, and to defeat our autocratic adversaries.Anthropic has therefore worked proactively to deploy our models to the Department of War and the intelligence community. We were the first frontier AI company to deploy our models in the US government’s classified networks, the first to deploy them at the National Laboratories, and the first to provide custom models for national security customers
AI PolicyAI GovernanceNational SecurityAI Ethics
Research arXiv (Artificial Intelligence) Feb 27

LLM Novice Uplift on Dual-Use, In Silico Biology Tasks

By Chen Bo Calvin Zhang, Christina Q. Knight, Nicholas Kruus, Jason Hausenloy, Pedro Medeiros, Nathaniel Li, Aiden Kim, Yury Orlovskiy, Coleman Breen, Bryce Cai, Jasper G\"otting, Andrew Bo Liu, Samira Nedungadi, Paula Rodriguez, Yannis Yiming He, Mohamed Shaaban, Zifan Wang, Seth Donoughe, Julian Michael

75 score
AI Analysis

Conducts a multi-model human uplift study showing LLM access makes novices 4.16x more accurate on biosecurity-relevant biology tasks compared to internet-only access. On some benchmarks, novices with LLMs matched or exceeded domain experts.

arXiv:2602.23329v1 Announce Type: new Abstract: Large language models (LLMs) perform increasingly well on biology benchmarks, but it remains unclear whether they uplift novice users -- i.e., enable humans to perform better than with internet-only resources. This uncertainty is central to understanding both scientific acceleration and dual-use risk. We conducted a multi-model, multi-benchmark human uplift study comparing novices with LLM access versus internet-only access across eight biosecurit
AI SafetyBiosecurityDual-Use RiskHuman-AI Collaboration
Research arXiv (Artificial Intelligence) Feb 27

ArchAgent: Agentic AI-driven Computer Architecture Discovery

By Raghav Gupta, Akanksha Jain, Abraham Gonzalez, Alexander Novikov, Po-Sen Huang, Matej Balog, Marvin Eisenberger, Sergey Shirobokov, Ng\^an V\~u, Martin Dixon, Borivoje Nikoli\'c, Parthasarathy Ranganathan, Sagar Karandikar

72 score
AI Analysis

Presents ArchAgent, built on AlphaEvolve, that automatically discovers computer architecture designs—specifically cache replacement policies. In two days without human intervention, it generated policies competitive with state-of-the-art within an established design competition framework.

arXiv:2602.22425v1 Announce Type: new Abstract: Agile hardware design flows are a critically needed force multiplier to meet the exploding demand for compute. Recently, agentic generative AI systems have demonstrated significant advances in algorithm design, improving code efficiency, and enabling discovery across scientific domains. Bridging these worlds, we present ArchAgent, an automated computer architecture discovery system built on AlphaEvolve. We show ArchAgent's ability to automatical
AI for HardwareAutomated DiscoveryAgentic AI
Research arXiv (Artificial Intelligence) Feb 27

Training Agents to Self-Report Misbehavior

By Bruce W. Lee, Chen Yueh-Han, Tomek Korbak

72 score
AI Analysis

As covered in Research yesterday, Proposes 'self-incrimination training' that trains AI agents (GPT-4.1 and Gemini-2.0) to call a report_scheming() tool when behaving deceptively. Significantly reduces undetected attack rates while preserving instruction hierarchy, outperforming monitors and alignment baselines.

arXiv:2602.22303v1 Announce Type: cross Abstract: Frontier AI agents may pursue hidden goals while concealing their pursuit from oversight. Alignment training aims to prevent such behavior by reinforcing the correct goals, but alignment may not always succeed and can lead to unwanted side effects. We propose self-incrimination training, which instead trains agents to produce a visible signal when they covertly misbehave. We train GPT-4.1 and Gemini-2.0 agents to call a report_scheming() tool wh
AI SafetyAlignmentDeceptive AIAgent Safety
Research arXiv (Machine Learning) Feb 27

Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA

By Hai Huang, Yann LeCun, Randall Balestriero

72 score
AI Analysis

Introduces Semantic Tube Prediction, a JEPA-style regularizer based on the Geodesic Hypothesis that token sequences trace geodesics on a semantic manifold, improving LLM data efficiency by 2-3x and challenging standard scaling laws.

arXiv:2602.22617v1 Announce Type: new Abstract: Large Language Models (LLMs) obey consistent scaling laws -- empirical power-law fits that predict how loss decreases with compute, data, and parameters. While predictive, these laws are descriptive rather than prescriptive: they characterize typical training, not optimal training. Surprisingly few works have successfully challenged the data-efficiency bounds implied by these laws -- which is our primary focus. To that end, we introduce the Geodes
Language ModelsScaling LawsJEPATraining Efficiency
Research arXiv (Computation and Language) Feb 27

AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

By Abhay Sheshadri, Aidan Ewart, Kai Fronsdal, Isha Gupta, Samuel R. Bowman, Sara Price, Samuel Marks, Rowan Wang

72 score
AI Analysis

AuditBench introduces a benchmark of 56 language models with implanted hidden behaviors (sycophancy, opposition to AI regulation, secret loyalties) that models don't confess when directly asked, used to evaluate alignment auditing techniques with an investigator agent.

arXiv:2602.22755v1 Announce Type: new Abstract: We introduce AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors. Each model has one of 14 concerning behaviors--such as sycophantic deference, opposition to AI regulation, or secret geopolitical loyalties--which it does not confess to when directly asked. AuditBench models are highly diverse--some are subtle, while others are overt, and we use varying training techniques both for
AI SafetyAlignmentBenchmarksDeception Detection
Research LessWrong Feb 26

How eval awareness might emerge in training

By Igor Ivanov

72 score
AI Analysis

Explores how evaluation awareness might emerge during model training rather than being purely a capability scaling phenomenon. Notes Apollo Research found Claude Opus 4.6 had high eval awareness that undermined safety testing.

Intro This post explores which aspects of model training lead to eval awareness and how it might help us mitigate it.The question is urgent. When Apollo Research conducted pre-deployment testing of Claude Opus 4.6, they reported:Apollo Research was given access to an early checkpoint of Claude Opus 4.6 on January 24th and an additional checkpoint on January 26th. During preliminary testing, Apollo did not find any instances of egregious misalignment, but observed high levels of verbalized e
AI SafetyEvaluationAlignmentEval Awareness
Research arXiv (Artificial Intelligence) Feb 27

Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?

By Pengxiang Li, Dilxat Muhtar, Lu Yin, Tianlong Chen, Shiwei Liu

70 score
AI Analysis

Analyzes why diffusion language models converge to autoregressive-like decoding despite being designed for parallel generation, attributing this to a mismatch between DLM objectives and sequential training data structure. Proposes NAP for truly non-autoregressive generation.

arXiv:2602.23225v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) are often advertised as enabling parallel token generation, yet practical fast DLMs frequently converge to left-to-right, autoregressive (AR)-like decoding dynamics. In contrast, genuinely non-AR generation is promising because it removes AR's sequential bottleneck, better exploiting parallel hardware to reduce synchronization/communication overhead and improve latency scaling with output length. We argue that a
Diffusion ModelsLanguage ModelsNon-Autoregressive GenerationArchitecture
Research arXiv (Artificial Intelligence) Feb 27

A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring

By Usman Anwar, Julianna Piskorz, David D. Baek, David Africa, Jim Weatherall, Max Tegmark, Christian Schroeder de Witt, Mihaela van der Schaar, David Krueger

68 score
AI Analysis

Proposes a decision-theoretic framework for detecting steganography in LLM outputs, addressing the challenge that classical steganography detection requires a known reference distribution. Central insight: steganography creates information asymmetry between decoders and non-decoders.

arXiv:2602.23163v1 Announce Type: new Abstract: Large language models are beginning to show steganographic capabilities. Such capabilities could allow misaligned models to evade oversight mechanisms. Yet principled methods to detect and quantify such behaviours are lacking. Classical definitions of steganography, and detection methods based on them, require a known reference distribution of non-steganographic signals. For the case of steganographic reasoning in LLMs, knowing such a reference di
AI SafetySteganographyLanguage ModelsAlignment
Research arXiv (Artificial Intelligence) Feb 27

Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability

By Bum Jun Kim, Shohei Taniguchi, Makoto Kawano, Yusuke Iwasawa, Yutaka Matsuo

68 score
AI Analysis

RKSP uses Koopman spectral analysis on a single forward pass at initialization to predict transformer training divergence with 0.995 AUROC, providing failure probability estimates before expensive training runs begin.

arXiv:2602.22988v1 Announce Type: cross Abstract: Training divergence in transformers wastes compute, yet practitioners discover instability only after expensive runs begin. They therefore need an expected probability of failure for a transformer before training starts. Our study of Residual Koopman Spectral Profiling (RKSP) provides such an estimate. From a single forward pass at initialization, RKSP extracts Koopman spectral features by applying whitened dynamic mode decomposition to layer-wi
Training StabilityTransformersEfficiencyTheory
Research arXiv (Artificial Intelligence) Feb 27

Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs

By Jayadev Billa

68 score
AI Analysis

This paper formalizes 'modality collapse' in multimodal LLMs as a mismatched decoder problem, showing that modality-specific information (speaker identity, emotion) survives through LLM layers but the text-trained decoder cannot use it. Bounded by Generalized Mutual Information.

arXiv:2602.23136v1 Announce Type: cross Abstract: Multimodal LLMs can process speech and images, but they cannot hear a speaker's voice or see an object's texture. We show this is not a failure of encoding: speaker identity, emotion, and visual attributes survive through every LLM layer (3--55$\times$ above chance in linear probes), yet removing 64--71% of modality-specific variance improves decoder loss. The decoder has no learned use for these directions; their presence is noise. We formali
Multimodal ModelsInformation TheoryLanguage ModelsTheory