Category intelligence

Research Briefing — February 20, 2026

387 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is headlined by a striking privacy result and several paradigm-challenging contributions to language modeling, safety, and interpretability.

In safety and alignment, fail-closed alignment identifies that current LLM safety is structurally fragile—removing a single refusal feature collapses alignment entirely. DeepMind applies AlphaEvolve to automatically discover novel multiagent learning algorithms. A systematic study of 60 LLM benchmarks characterizes saturation dynamics and what drives benchmark obsolescence.

Key Themes

AI Safety, Privacy & Oversight · 8AI Safety & Alignment · 30Language Models & Alignment · 10LLM Training & Optimization · 10Agent Systems & Tool Use · 15LLM-Driven Algorithm Discovery · 3Mechanistic Interpretability · 9LLM Reasoning · 6Coding Agents and Software Engineering · 1Benchmarking & Evaluation · 12

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Feb 20

Large-scale online deanonymization with LLMs

By Simon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni, Nicholas Carlini, Florian Tram\`er

88 score
AI Analysis

Demonstrates that LLMs can perform large-scale deanonymization of online users, re-identifying Hacker News users and Anthropic interview participants from pseudonymous profiles. Implements a scalable pipeline using feature extraction, semantic embeddings, and reasoning for matching across databases.

arXiv:2602.16800v1 Announce Type: cross Abstract: We show that large language models can be used to perform at-scale deanonymization. With full Internet access, our agent can re-identify Hacker News users and Anthropic Interviewer participants at high precision, given pseudonymous online profiles and conversations alone, matching what would take hours for a dedicated human investigator. We then design attacks for the closed-world setting. Given two databases of pseudonymous individuals, each co
AI SafetyPrivacyLanguage ModelsSecurity
Research arXiv (Artificial Intelligence) Feb 20

One-step Language Modeling via Continuous Denoising

By Chanhyuk Lee, Jaehoon Yoo, Manan Agarwal, Sheel Shah, Jerry Huang, Aditi Raghunathan, Seunghoon Hong, Nicholas M. Boffi, Jinwoo Kim

78 score
AI Analysis

Proposes flow-based language models (FLM) that perform Euclidean denoising over one-hot token encodings, outperforming discrete diffusion models in both quality and speed. Demonstrates strong performance in the few-step generation regime where discrete diffusion degrades sharply.

arXiv:2602.16813v1 Announce Type: cross Abstract: Language models based on discrete diffusion have attracted widespread interest for their potential to provide faster generation than autoregressive models. In practice, however, they exhibit a sharp degradation of sample quality in the few-step regime, failing to realize this promise. Here we show that language models leveraging flow-based continuous denoising can outperform discrete diffusion in both quality and speed. By revisiting the fundame
Language ModelsDiffusion ModelsNon-Autoregressive GenerationNovel Architectures
Research arXiv (Artificial Intelligence) Feb 20

Xray-Visual Models: Scaling Vision models on Industry Scale Data

By Shlok Mishra, Tsung-Yu Lin, Linda Wang, Hongli Xu, Yimin Liu, Michael Hsu, Chaitanya Ahuja, Hao Yuan, Jianpeng Cheng, Hong-You Chen, Haoyuan Xu, Chao Li, Abhijeet Awasthi, Jihye Moon, Don Husa, Michael Ge, Sumedha Singla, Arkabandhu Chowdhury, Phong Dingh, Satya Narayan Shukla, Yonghuan Yang, David Jacobs, Qi Guo, Jun Xiao, Xiangjun Fan, Aashu Singh

75 score
AI Analysis

Xray-Visual is a unified vision model trained on 15B+ image-text pairs and 10B video-hashtag pairs from Facebook/Instagram. Uses a three-stage training pipeline combining MAE, hashtag classification, and CLIP-style contrastive learning.

arXiv:2602.16918v1 Announce Type: cross Abstract: We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 billion curated image-text pairs and 10 billion video-hashtag pairs from Facebook and Instagram, employing robust data curation pipelines that incorporate balancing and noise suppression strategies to maximize semantic diversity while minimizing label noise. We introduc
Computer VisionFoundation ModelsVision-Language ModelsIndustry Scale
Research arXiv (Machine Learning) Feb 20

Escaping the Cognitive Well: Efficient Competition Math with Off-the-Shelf Models

By Xingyu Dang, Rohit Agarwal, Rodrigo Porto, Anirudh Goyal, Liam H Fowl, Sanjeev Arora

75 score
AI Analysis

Presents an inference pipeline achieving top performance on IMO-style math problems at orders of magnitude lower cost than competing methods, using only off-the-shelf models. Identifies the 'Cognitive Well' problem where solver-grader pipelines converge to wrong solutions.

arXiv:2602.16793v1 Announce Type: new Abstract: In the past year, custom and unreleased math reasoning models reached gold medal performance on the International Mathematical Olympiad (IMO). Similar performance was then reported using large-scale inference on publicly available models but at prohibitive costs (e.g., 3000 USD per problem). In this work, we present an inference pipeline that attains best-in-class performance on IMO-style math problems at an average inference cost orders of magnit
Mathematical ReasoningLLM ReasoningInference Efficiency
Research arXiv (Artificial Intelligence) Feb 20

Discovering Multiagent Learning Algorithms with Large Language Models

By Zun Li, John Schultz, Daniel Hennes, Marc Lanctot

74 score
AI Analysis

Uses AlphaEvolve (LLM-driven evolutionary search) to automatically discover new multiagent learning algorithms, evolving novel variants of Counterfactual Regret Minimization and Policy Space Response Oracles for imperfect-information games.

arXiv:2602.16928v1 Announce Type: cross Abstract: Much of the advancement of Multi-Agent Reinforcement Learning (MARL) in imperfect-information games has historically depended on manual iterative refinement of baselines. While foundational families like Counterfactual Regret Minimization (CFR) and Policy Space Response Oracles (PSRO) rest on solid theoretical ground, the design of their most effective variants often relies on human intuition to navigate a vast algorithmic design space. In this
Multi-Agent Reinforcement LearningAlgorithm DiscoveryLLM-Driven EvolutionGame Theory
Research arXiv (Machine Learning) Feb 20

Fail-Closed Alignment for Large Language Models

By Zachary Coalson, Beth Sohler, Aiden Gabriel, Sanghyun Hong

74 score
AI Analysis

Identifies that current LLM alignment is 'fail-open' - suppressing a single dominant refusal feature causes alignment collapse. Proposes fail-closed alignment that builds redundant, independent refusal pathways that survive partial failures.

arXiv:2602.16977v1 Announce Type: new Abstract: We identify a structural weakness in current large language model (LLM) alignment: modern refusal mechanisms are fail-open. While existing approaches encode refusal behaviors across multiple latent features, suppressing a single dominant feature$-$via prompt-based jailbreaks$-$can cause alignment to collapse, leading to unsafe generation. Motivated by this, we propose fail-closed alignment as a design principle for robust LLM safety: refusal mecha
AI SafetyAlignmentLLM Robustness
Research arXiv (Artificial Intelligence) Feb 20

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

By Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vil\'em Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek \v{S}uppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Pawar, Robert Scholz, Arjun Subramonian, Jingwei Ni, Mykel Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, Irene Solaiman

72 score
AI Analysis

Systematically studies benchmark saturation across 60 LLM benchmarks from major model developers, analyzing 14 properties to identify factors driving saturation. Finds that nearly half of benchmarks are saturated and identifies design properties that predict saturation rates.

arXiv:2602.16763v1 Announce Type: new Abstract: Artificial Intelligence (AI) benchmarks play a central role in measuring progress in model development and guiding deployment decisions. However, many benchmarks quickly become saturated, meaning that they can no longer differentiate between the best-performing models, diminishing their long-term value. In this study, we analyze benchmark saturation across 60 Large Language Model (LLM) benchmarks selected from technical reports by major model deve
BenchmarkingLanguage ModelsAI Evaluation
Research arXiv (Artificial Intelligence) Feb 20

References Improve LLM Alignment in Non-Verifiable Domains

By Kejian Shi, Yixin Liu, Peifeng Wang, Alexander R. Fabbri, Shafiq Joty, Arman Cohan

72 score
AI Analysis

Investigates using reference outputs to improve LLM-based evaluators for alignment in non-verifiable domains where RLVR cannot be directly applied. Shows that reference-guided evaluation improves weaker judges and enables better RLHF training.

arXiv:2602.16802v1 Announce Type: cross Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) has shown strong effectiveness in reasoning tasks, it cannot be directly applied to non-verifiable domains lacking ground-truth verifiers, such as LLM alignment. In this work, we investigate whether reference-guided LLM-evaluators can bridge this gap by serving as soft "verifiers". First, we design evaluation protocols that enhance LLM-based evaluators for LLM alignment using reference
AlignmentRLHFLanguage ModelsEvaluation
Research arXiv (Artificial Intelligence) Feb 20

The Anxiety of Influence: Bloom Filters in Transformer Attention Heads

By Peter Balogh

72 score
AI Analysis

Identifies transformer attention heads that function as membership testers ("has this token appeared before?") and shows they follow Bloom filter theory. Demonstrates these heads form a spectrum of strategies across GPT-2 and Pythia models, with some exceeding classical Bloom filter capacity bounds.

arXiv:2602.17526v1 Announce Type: cross Abstract: Some transformer attention heads appear to function as membership testers, dedicating themselves to answering the question "has this token appeared before in the context?" We identify these heads across four language models (GPT-2 small, medium, and large; Pythia-160M) and show that they form a spectrum of membership-testing strategies. Two heads (L0H1 and L0H5 in GPT-2 small) function as high-precision membership filters with false positive rat
Mechanistic InterpretabilityTransformer ArchitectureLanguage Models
Research arXiv (Artificial Intelligence) Feb 20

Towards Anytime-Valid Statistical Watermarking

By Baihe Huang, Eric Xu, Kannan Ramchandran, Jiantao Jiao, Michael I. Jordan

72 score
AI Analysis

Develops the first e-value-based watermarking framework for LLMs that enables anytime-valid inference with principled early stopping, addressing the limitation of fixed-horizon hypothesis testing in existing watermarking methods.

arXiv:2602.17608v1 Announce Type: cross Abstract: The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has emerged as a promising solution, existing methods suffer from two critical limitations: the lack of a principled approach for selecting sampling distributions and the reliance on fixed-horizon hypothesis testing, which precludes valid early stopping. In this paper, we bri
AI SafetyLLM WatermarkingStatistical Testing
Research arXiv (Machine Learning) Feb 20

Formal Mechanistic Interpretability: Automated Circuit Discovery with Provable Guarantees

By Itamar Hadad, Guy Katz, Shahaf Bassan

72 score
AI Analysis

Proposes formal automated circuit discovery for neural networks with provable guarantees using neural network verification. Provides three types of guarantees: input domain robustness, robust patching, and faithful compression.

arXiv:2602.16823v1 Announce Type: new Abstract: *Automated circuit discovery* is a central tool in mechanistic interpretability for identifying the internal components of neural networks responsible for specific behaviors. While prior methods have made significant progress, they typically depend on heuristics or approximations and do not offer provable guarantees over continuous input domains for the resulting circuits. In this work, we leverage recent advances in neural network verification to
Mechanistic InterpretabilityFormal VerificationAI Safety
Research arXiv (Machine Learning) Feb 20

Unified Latents (UL): How to train your latents

By Jonathan Heek, Emiel Hoogeboom, Thomas Mensink, Tim Salimans

72 score
AI Analysis

Presents Unified Latents (UL), a framework for jointly learning latent representations with diffusion prior regularization and diffusion decoding. Achieves FID 1.4 on ImageNet-512 and state-of-the-art FVD 1.3 on Kinetics-600.

arXiv:2602.17270v1 Announce Type: new Abstract: We present Unified Latents (UL), a framework for learning latent representations that are jointly regularized by a diffusion prior and decoded by a diffusion model. By linking the encoder's output noise to the prior's minimum noise level, we obtain a simple training objective that provides a tight upper bound on the latent bitrate. On ImageNet-512, our approach achieves competitive FID of 1.4, with high reconstruction quality (PSNR) while requirin
Generative ModelsImage GenerationVideo Generation