Category intelligence

Research Briefing — May 25, 2026

446 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by safety/red-teaming findings on frontier models and foundational advances in training, retrieval, and RL.

Safety, Alignment & Red-Teaming:

Foundations & Systems:

RL & Evaluation Methodology:

Key Themes

AI Safety & Evaluation · 6AI Safety, Robustness & Red-teaming · 9Reinforcement Learning & Post-Training · 10AI Safety & Red-Teaming · 6Frontier LLM Evaluation & Benchmarks · 9Agentic AI Systems · 12LLM Reasoning & Interpretability · 6Mechanistic Interpretability & Steering · 7AI Agents & Workflow Automation · 6LLM Inference Efficiency · 6

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) May 25

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

By Shubham Agarwal, Alexander Krentsel, Shu Liu, Mert Cemri, Audrey Cheng, Rui Meng, Tomas Pfister, Chun-Liang Li, Sylvia Ratnasamy, Aditya Parameswaran, Matei Zaharia, Ion Stoica, Mohsen Lesani

78 score
AI Analysis

Inductive Deductive Synthesis (IDS) jointly synthesizes implementation and proof for formally verified distributed systems, where SOTA agents (Codex/GPT-5.4, Claude Opus 4.6) succeed on only 2/7 tasks. Strong author list and significant capability gap addressed.

AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage that testing alone cannot provide. Distributed systems are a prime example: properties such as consistency between reads and writes must hold under every possible interleaving of events. Mechanized formal verification can guarantee such correctness, but typically demands months to years of expert effort. As evidence, even SOTA coding agents (Codex
AI AgentsFormal VerificationCode Generation
Research arXiv (Machine Learning) May 25

Test-Time Training Undermines Safety Guardrails

By Simone Antonelli, Sadegh Akhondzadeh, Aleksandar Bojchevski

75 score
AI Analysis

Demonstrates that Test-Time Training enables new jailbreak attacks with 95% Attack Success Rate over 10 trials under LoRA, transferring across model families. Important safety finding for adaptive inference.

Test-Time Training (TTT) is an emerging paradigm that enables models to adapt their parameters during inference, improving performance on tasks such as few-shot learning, retrieval-augmented generation, and complex reasoning. However, this dynamic adaptation introduces new vulnerabilities that adversaries can exploit to jailbreak models. We identify three threat models for TTT and demonstrate how attackers can leverage them to bypass safety filters. Our results show that TTT can significantly in
AI SafetyJailbreakingTest-Time Training
Research arXiv (Machine Learning) May 25

Decomposing and Measuring Evaluation Awareness

By Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Sahar Abdelnabi, Maksym Andriushchenko

75 score
AI Analysis

Decomposes evaluation awareness into environment recognizability and model propensity, operationalizing through 8 trigger factors and CoT monitoring across 9 frontier models and 4 benchmarks. Important framework for studying evaluation gaming.

Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a shared foundation, conflating properties of the evaluation with properties of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component (how recognizable the task is) and a model component that separates recognition from propensi
AI SafetyEvaluation AwarenessAlignment
Research arXiv (Machine Learning) May 25

Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization

By Alexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Yan Zuo, Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi, Sameera Ramasinghe

72 score
AI Analysis

UPMs introduce time-varying invertible transforms at participant boundaries in distributed model training so that no participant ever holds extractable weights, while preserving network function. Tested on Qwen-2.5 and Llama-3.2 with negligible perplexity loss.

We consider a decentralized setup in which the participants collaboratively train and serve a large neural network, and where each participant only processes a subset of the model. In this setup, we explore the possibility of unmaterializable weights, where a full weight set is never available to any one participant. We introduce Unextractable Protocol Models (UPMs): a training and inference framework that leverages the sharded model setup to ensure model shards (i.e., subsets) held by participa
Distributed TrainingPrivacyModel Security
Research arXiv (Machine Learning) May 25

Goal-Conditioned Agents that Learn Everything All at Once

By Michael Matthews, Matthew Jackson, Michael Beukman, Thomas Foster, Alistair Letcher, Scott Fujimoto, C\'edric Colas, Jakob Foerster

72 score
AI Analysis

LEO enables efficient all-goals off-policy learning by jointly outputting values and actions for every goal in a single forward pass, avoiding naive relabelling cost. Significantly outperforms baselines on goal-conditioned Craftax. Authors include Jakob Foerster and others from established RL groups.

A goal-conditioned reinforcement learning agent exploring an environment will see a wealth of information throughout a trajectory, most of which is discarded when only performing on-policy updates with respect to the commanded goal. All-goals learning, where each transition is used for learning off-policy with respect to every goal, allows agents to extract maximal information, however it is usually computationally infeasible when done via naive relabelling. This can be overcome by jointly outpu
Reinforcement LearningGoal-Conditioned RLOff-Policy Learning
72 score
AI Analysis

Empirical study showing that geopolitical bias in LLMs originates from post-training (not pretraining), with developer-aligned shifts after instruction tuning. Strongest in Qwen 2.5 (China-favorability shifts dramatically). Important finding for alignment research.

It has generally been assumed that geopolitical bias in language models originates from the training data used during the pre-training phase. We tested seven open-weight LLM pairs consisting of the base model (pre-training only) and the chat model (pre-training and post-training) from seven labs on a paired-scenario forced-choice probe over 28 country pairs in English, French, and Chinese, and found that geopolitical bias originates in post-training rather than in pre-training. Across seven AI l
AI SafetyBiasAlignmentPost-Training
Research arXiv (Computer Vision) May 25

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

By Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou

70 score
AI Analysis

Investigates whether VLM benchmarks actually test visual understanding by showing that removing image tokens barely degrades performance on hallucination benchmarks. Reveals significant disconnect between benchmark accuracy and grounded visual reasoning.

Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly reflect reliance on visual evidence. Motivated by a surprising observation that removing a substantial fraction of image tokens only degrades model performance very slightly on a widely used hallucination benchmark, we systematically investigate this mismatch in a set of open-source VLMs. Our analysis spans multiple leve
Vision-Language ModelsEvaluationBenchmarks
Research arXiv (Computation and Language) May 25

Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

By Dongxin Guo, Jikun Wu, Siu Ming Yiu

70 score
AI Analysis

Uses sparse autoencoders to decompose GPT-2 XL and Llama-3.1-8B and shows semantic features alone recover 94% of brain encoding performance, mapping LLM features onto cortical semantic topography. Strong mechanistic evidence linking SAE features to neuroscience.

Intermediate layers of large language models (LLMs) best predict human brain responses to language, one of the most robust findings in computational neurolinguistics, yet why remains mechanistically unexplained. We address this gap by bridging sparse autoencoders (SAEs) from mechanistic interpretability with neural encoding models, decomposing GPT-2 XL and Llama-3.1-8B into 16K-32K interpretable features per layer. A human-validated taxonomy ($\kappa \geq 0.74$) reveals that semantic features al
InterpretabilitySparse AutoencodersNeuroscience
Research arXiv (Computation and Language) May 25

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

By Casey Ford, Madison Van Doren, Sicheng Jin, and Emily Dix

70 score
AI Analysis

Cross-lingual multimodal red-teaming of Claude Sonnet 4.5, GPT-5, Pixtral Large, Qwen Omni in en-US vs es-MX, finding language non-uniformly shifts vulnerability surface (52K ratings, 9 annotators per group).

The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures. We present the first systematic cross-lingual, multimodal red-teaming study comparing jailbreak vulnerability in US English (en-US) and Mexican Spanish (es-MX) across four frontier MLLMs: Claude Sonnet 4.5, GPT-5, Pixtral Large, and Qwen Omni. Using a fixed adversarial benchmark of 363 diverse prompt scenarios administered in text-only and mult
AI SafetyJailbreakingMultilingualRed-teaming
Research arXiv (cs.CR) May 25

PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs

By Luze Sun, Anshuman Suri, Harsh Chaudhari, Cristina Nita-Rotaru, Alina Oprea

70 score
AI Analysis

PoisonForge: 4-dimensional task-level poisoning benchmark for instruction tuning showing 11/12 models compromised at 1% poison budget. Strong adversarial ML team (Oprea, Nita-Rotaru).

When practitioners fine-tune LLMs on unvetted datasets, an adversary can exploit the data supply chain through task-level poisoning: inserting a small number of crafted instruction-response pairs that cause the model to embed attacker-specified entities, such as a country, in outputs for a targeted task family while behaving normally elsewhere. We introduce PoisonForge, a benchmark that parameterizes this threat along four dimensions (bias type, poisoning mode, appearance count, and target outpu
AI SafetyData PoisoningInstruction Tuning
Research arXiv (Machine Learning) May 25

FastKernels: Benchmarking GPU Kernel Generation in Production

By Gabriele Oliaro, Yichao Fu, May Jiang, Owen Lu, Junli Wang, Zhihao Jia, Hao Zhang, and Samyam Rajbhandari

70 score
AI Analysis

FastKernels: production-aligned benchmark of 46 architectures for LLM-agent GPU kernel generation, exposing how existing benchmarks reward sandbox tricks rather than real systems performance. Includes Zhihao Jia, Samyam Rajbhandari.

LLM-based agents for GPU kernel generation are advancing rapidly, yet their progress is fundamentally constrained by the benchmarks they optimize against. Existing benchmarks are poorly aligned with production inference frameworks: they evaluate kernels on a single GPU with synthetic inputs, ignore the surrounding compilation stack, and reward replicating known optimizations rather than discovering new ones. The resulting reward signals are misleading: agents learn to generate kernels that score
AI for CodeGPU OptimizationBenchmarks
Research arXiv (cs.CR) May 25

Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks

By Vivek Dahiya, Sunny Nehra, Vipul Dholariya, Bhavik Shangari, Chandra Khatri

70 score
AI Analysis

Dual-mode benchmark (white-box vulnerability detection + black-box web app pentesting) finding all six frontier models (GPT-5.4, Codex-5.3, Opus 4.6, Sonnet 4.6, Gemini 3.1 Pro, Gemini 3 Flash) produce 10-50% FP rates and only 4-19% GT coverage—argues for vertical foundation models.

We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detection (VulnLLM-R, across C/Java/Python) and black-box web application security testing (five production-style applications with 118 ground-truth vulnerabilities across 20+ CWE families, which we will open-source). We test six frontier models (GPT-5.4, Codex~5.3, Claude Opus~4.6, Sonnet~4.6, Gemini~3.1~Pro and Gemini~3~Flash) and two domain-specialized models acr
CybersecurityLLM EvaluationFrontier Models