Category intelligence

Research Briefing — January 7, 2026

358 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research centers on AI safety vulnerabilities and foundational model capabilities. A systematic study demonstrates extraction of copyrighted books from production LLMs using Best-of-N jailbreaking, raising significant legal implications. Critical stress-testing of Anthropic's SAE features reveals fragility in steering interventions, questioning interpretability claims.

Practical contributions include agent-permissions.json for web agent governance and a striking one-shot RL finding showing single-sample training produces improvements rivaling full datasets. Jacob Steinhardt's Oversight Assistants framework addresses scalable human oversight of AI systems.

Key Themes

Foundation Models · 3AI Safety & Governance · 5AI Safety and Security · 8Mechanistic Interpretability · 8LLM Safety & Alignment · 9LLM Applications & Agents · 18LLM Reasoning & Efficiency · 12AI Safety & Alignment · 24LLM Efficiency & Architecture · 6Benchmarking & Evaluation · 9

Primary evidence

Top Ranked Signals

Research arXiv (Computer Vision) Jan 7

NitroGen: An Open Foundation Model for Generalist Gaming Agents

By Lo\"ic Magne, Anas Awadalla, Guanzhi Wang, Yinzhen Xu, Joshua Belofsky, Fengyuan Hu, Joohwan Kim, Ludwig Schmidt, Georgia Gkioxari, Jan Kautz, Yisong Yue, Yejin Choi, Yuke Zhu, Linxi "Jim" Fan

82 score
AI Analysis
NitroGen is vision-action foundation model trained on 40,000 hours of gameplay across 1,000+ games. Demonstrates cross-game generalization with up to 52% improvement on unseen games.
We introduce NitroGen, a vision-action foundation model for generalist gaming agents that is trained on 40,000 hours of gameplay videos across more than 1,000 games. We incorporate three key ingredients: 1) an internet-scale video-action dataset constructed by automatically extracting player actions from publicly available gameplay videos, 2) a multi-game benchmark environment that can measure cross-game generalization, and 3) a unified vision-action model trained with large-scale behavior cloni
Foundation ModelsGame AIBehavior CloningVision-Action Models
Research arXiv (Computation and Language) Jan 7

Extracting books from production language models

By Ahmed Ahmed and A. Feder Cooper and Sanmi Koyejo and Percy Liang

82 score
AI Analysis
Investigates extraction of copyrighted books from production LLMs using two-phase procedure with Best-of-N jailbreak and iterative prompts. Demonstrates substantial memorization in deployed systems.
Many unresolved legal questions over LLMs and copyright center on memorization: whether specific training data have been encoded in the model's weights during training, and whether those memorized data can be extracted in the model's outputs. While many believe that LLMs do not memorize much of their training data, recent work shows that substantial amounts of copyrighted text can be extracted from open-weight models. However, it remains an open question if similar extraction is feasible for pro
AI SafetyCopyrightMemorizationLanguage Models
Research arXiv (cs.CY) Jan 7

Permission Manifests for Web Agents

By Samuele Marro, Alan Chan, Xinxing Ren, Lewis Hammond, Jesse Wright, Gurjyot Wanga, Tiziano Piccardi, Nuno Campos, Tobin South, Jialin Yu, Alex Pentland, Philip Torr, Jiaxin Pei

78 score
AI Analysis
Proposes agent-permissions.json, a robots.txt-style manifest for websites to specify allowed/disallowed interactions with LLM web agents. Addresses governance gap for sophisticated AI agents.
The rise of Large Language Model (LLM)-based web agents represents a significant shift in automated interactions with the web. Unlike traditional crawlers that follow simple conventions, such as robots.txt, modern agents engage with websites in sophisticated ways: navigating complex interfaces, extracting structured information, and completing end-to-end tasks. Existing governance mechanisms were not designed for these capabilities. Without a way to specify what interactions are and are not allo
AI GovernanceWeb AgentsAI Safety
Research arXiv (Robotics) Jan 7

InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation

By Junhao Cai, Zetao Cai, Jiafei Cao, Yilun Chen, Zeyu He, Lei Jiang, Hang Li, Hengjie Li, Yang Li, Yufei Liu, Yanan Lu, Qi Lv, Haoxiang Ma, Jiangmiao Pang, Yu Qiao, Zherui Qiu, Yanqing Shen, Xu Shi, Yang Tian, Bolun Wang, Hanqing Wang, Jiaheng Wang, Tai Wang, Xueyuan Wei, Chao Wu, Yiman Xie, Boyang Xing, Yuqiang Yang, Yuyin Yang, Qiaojun Yu, Feng Yuan, Jia Zeng, Jingjing Zhang, Shenghan Zhang, Shi Zhang, Zhuoma Zhaxi, Bowen Zhou, Yuanzhen Zhou, Yunsong Zhou, Hongrui Zhu, Yangkun Zhu, Yuchen Zhu

78 score
AI Analysis
Introduces InternVLA-A1, a unified Vision-Language-Action model using Mixture-of-Transformers architecture coordinating three experts for scene understanding, visual foresight generation, and action execution for robotic manipulation.
Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, but they inherently lack the capability to deduce physical world dynamics. Consequently, recent approaches have shifted toward World Models, typically formulated via video prediction; however, these methods often suffer from a lack of semantic grounding and exhibit brittleness when handling prediction errors. To synergi
RoboticsVision-Language-ActionWorld ModelsEmbodied AI
Research arXiv (Artificial Intelligence) Jan 7

Logical Phase Transitions: Understanding Collapse in LLM Logical Reasoning

By Xinglang Zhang, Yunyao Zhang, ZeLiang Chen, Junqing Yu, Wei Yang, Zikai Song

78 score
AI Analysis
Discovers 'Logical Phase Transitions' where LLM reasoning collapses abruptly beyond critical complexity rather than degrading smoothly. Proposes Neuro-Symbolic Curriculum Tuning based on this insight.
Symbolic logical reasoning is a critical yet underexplored capability of large language models (LLMs), providing reliable and verifiable decision-making in high-stakes domains such as mathematical reasoning and legal judgment. In this study, we present a systematic analysis of logical reasoning under controlled increases in logical complexity, and reveal a previously unrecognized phenomenon, which we term Logical Phase Transitions: rather than degrading smoothly, logical reasoning performance re
LLM ReasoningAI SafetyNeuro-Symbolic AIFailure Modes
Research arXiv (Machine Learning) Jan 7

When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability

By Raphael Ronge, Markus Maier, Frederick Eberhardt

78 score
AI Analysis
Stress-tests Anthropic's mechanistic interpretability claims by replicating SAE feature extraction and steering with open-source SAEs for Llama 3.1, finding substantial fragility in feature steering with sensitivity to layer, magnitude, and context.
Recent work by Anthropic on Mechanistic interpretability claims to understand and control Large Language Models by extracting human-interpretable features from their neural activation patterns using sparse autoencoders (SAEs). If successful, this approach offers one of the most promising routes for human oversight in AI safety. We conduct an initial stress-test of these claims by replicating their main results with open-source SAEs for Llama 3.1. While we successfully reproduce basic feature ext
Mechanistic InterpretabilityAI SafetySparse Autoencoders
Research arXiv (cs.SE) Jan 7

ProSoftArena: Benchmarking Hierarchical Capabilities of Multimodal Agents in Professional Software Environments

By Jiaxin Ai, Yukang Feng, Fanrui Zhang, Jianwen Sun, Zizhen Li, Chuanhao Li, Yifan Chang, Wenxiao Wu, Ruoxi Wang, Mingliang Zhai, Kaipeng Zhang

77 score
AI Analysis
ProSoftArena benchmark with 436 tasks across 13 professional applications for evaluating multimodal agents. First capability hierarchy for agent use of professional software with execution-based evaluation.
Multimodal agents are making rapid progress on general computer-use tasks, yet existing benchmarks remain largely confined to browsers and basic desktop applications, falling short in professional software workflows that dominate real-world scientific and industrial practice. To close this gap, we introduce ProSoftArena, a benchmark and platform specifically for evaluating multimodal agents in professional software environments. We establish the first capability hierarchy tailored to agent use o
Multimodal AgentsBenchmarkingSoftware Engineering
Research arXiv (Machine Learning) Jan 7

From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence

By Marc Finzi, Shikai Qiu, Yiding Jiang, Pavel Izmailov, J. Zico Kolter, Andrew Gordon Wilson

76 score
AI Analysis
Proposes 'epiplexity' as new information measure for computationally bounded agents, addressing paradoxes in Shannon information theory. Shows information can be constructed through deterministic transformations when computation is bounded.
Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to existing data? Can the learnable content in data be evaluated without considering a downstream task? On these questions, Shannon information and Kolmogorov complexity come up nearly empty-handed, in part because they assume observers with unlimited computational capacity and fail to target the useful information content. In
Information TheoryMachine Learning TheoryFoundations
Research arXiv (Computation and Language) Jan 7

Window-based Membership Inference Attacks Against Fine-tuned Large Language Models

By Yuetian Chen, Yuntao Du, Kaiyuan Zhang, Ashish Kundu, Charles Fleming, Bruno Ribeiro, Ninghui Li

76 score
AI Analysis
Introduces WBC (Window-Based Comparison), a membership inference attack against fine-tuned LLMs using sliding windows with sign-based aggregation. Challenges the global-averaging paradigm by showing membership signals are localized.
Most membership inference attacks (MIAs) against Large Language Models (LLMs) rely on global signals, like average loss, to identify training data. This approach, however, dilutes the subtle, localized signals of memorization, reducing attack effectiveness. We challenge this global-averaging paradigm, positing that membership signals are more pronounced within localized contexts. We introduce WBC (Window-Based Comparison), which exploits this insight through a sliding window approach with sign-b
AI SecurityPrivacyMembership InferenceLLM Safety
Research arXiv (Machine Learning) Jan 7

Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control

By Harshvardhan Saini, Yiming Tang, Dianbo Liu

76 score
AI Analysis
Proposes using gradient ascent to discover prompts that align with persona directions identified through mechanistic interpretability. Bridges interpretability research with prompt engineering for persona control.
Controlling emergent behavioral personas (e.g., sycophancy, hallucination) in Large Language Models (LLMs) is critical for AI safety, yet remains a persistent challenge. Existing solutions face a dilemma: manual prompt engineering is intuitive but unscalable and imprecise, while automatic optimization methods are effective but operate as "black boxes" with no interpretable connection to model internals. We propose a novel framework that adapts gradient ascent to LLMs, enabling targeted prompt di
Mechanistic InterpretabilityPrompt EngineeringAI SafetyAlignment
Research arXiv (Machine Learning) Jan 7

One Sample to Rule Them All: Extreme Data Efficiency in RL Scaling

By Yiyuan Li, Zhen Huang, Yanan Wu, Weixun Wang, Xuefeng Li, Yijia Luo, Wenbo Su, Bo Zheng, Pengfei Liu

76 score
AI Analysis
Demonstrates remarkable effectiveness of one-shot learning in RL for LLMs - a single strategically selected math reasoning sample produces improvements across physics, chemistry, and biology through 'polymath learning' framework.
The reasoning ability of large language models (LLMs) can be unleashed with reinforcement learning (RL) (OpenAI, 2024; DeepSeek-AI et al., 2025a; Zeng et al., 2025). The success of existing RL attempts in LLMs usually relies on high-quality samples of thousands or beyond. In this paper, we challenge fundamental assumptions about data requirements in RL for LLMs by demonstrating the remarkable effectiveness of one-shot learning. Specifically, we introduce polymath learning, a framework for design
Reinforcement LearningData EfficiencyTransfer Learning
75 score
AI Analysis
Presents Chronicals, an LLM fine-tuning framework claiming 3.51x speedup over Unsloth through fused Triton kernels, Cut Cross-Entropy, LoRA+ differential learning rates, and sequence packing optimizations.
Large language model fine-tuning is bottlenecked by memory: a 7B parameter model requires 84GB--14GB for weights, 14GB for gradients, and 56GB for FP32 optimizer states--exceeding even A100-40GB capacity. We present Chronicals, an open-source training framework achieving 3.51x speedup over Unsloth through four synergistic optimizations: (1) fused Triton kernels eliminating 75% of memory traffic via RMSNorm (7x), SwiGLU (5x), and QK-RoPE (2.3x) fusion; (2) Cut Cross-Entropy reducing logit memory
LLM TrainingEfficiencySystems Optimization