Category intelligence

Research Briefing — February 18, 2026

304 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research centers on alignment fragility, scaling laws, and emergent risks in agentic systems.

  • The Geometry of Alignment Collapse proves safety alignment concentrates in brittle low-dimensional subspaces, explaining why fine-tuning breaks safety guarantees
  • Prescriptive Scaling Laws predict downstream task accuracy as a function of pre-training compute using smoothed quantile regression across 500+ tasks
  • The Obfuscation Atlas (Anthropic-affiliated) maps how deception naturally emerges when LLMs are trained against white-box monitors, introducing a taxonomy of obfuscation strategies
  • GLM-5 from Zhipu/Tsinghua presents a foundation model for agentic engineering with novel asynchronous RL infrastructure and agent-specific algorithms

On data and training stability, ÜberWeb reveals that multilingual regression degrades English performance at 20T-token scale, while STAPO identifies that just ~0.01% of spurious tokens drive late-stage RL training collapse. ResearchGym benchmarks AI agents on end-to-end research tasks, exposing a stark capability-reliability gap.

Key Themes

AI Safety and Alignment · 20Foundation Models and LLMs · 8Language Models · 15Scaling and Data Curation · 4AI Safety & Alignment · 10AI Agents · 12Interpretability · 6Interpretability & Representation Engineering · 7AI Safety, Alignment, and Verification · 7Steganography & CoT Monitoring · 1

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Feb 18

The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety

By Max Springer, Chung Peng Lee, Blossom Metevier, Jane Castleman, Bohdan Turbal, Hayoung Jung, Zeyu Shen, Aleksandra Korolova

83 score
AI Analysis

Proves that safety alignment in LLMs concentrates in low-dimensional subspaces with sharp curvature, creating brittle structure that gradient descent cannot detect or defend. Shows that the common explanation of orthogonality between fine-tuning and safety directions is structurally unstable.

arXiv:2602.15799v1 Announce Type: cross Abstract: Fine-tuning aligned language models on benign tasks unpredictably degrades safety guardrails, even when training data contains no harmful content and developers have no adversarial intent. We show that the prevailing explanation, that fine-tuning updates should be orthogonal to safety-critical directions in high-dimensional parameter space, offers false reassurance: we show this orthogonality is structurally unstable and collapses under the dyna
AI SafetyAlignmentFine-TuningTheory
Research arXiv (Artificial Intelligence) Feb 18

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

By Hanlin Zhang, Jikai Jin, Vasilis Syrgkanis, Sham Kakade

82 score
AI Analysis

Develops prescriptive scaling laws that predict downstream task accuracy as a function of pre-training compute using smoothed quantile regression on 5K+ model observations. Validates temporal reliability by fitting on earlier model generations and evaluating on later releases, finding stable capability boundaries across tasks.

arXiv:2602.15327v1 Announce Type: cross Abstract: For deploying foundation models, practitioners increasingly need prescriptive scaling laws: given a pre training compute budget, what downstream accuracy is attainable with contemporary post training practice, and how stable is that mapping as the field evolves? Using large scale observational evaluations with 5k observational and 2k newly sampled data on model performance, we estimate capability boundaries, high conditional quantiles of benchma
Scaling LawsLanguage ModelsFoundation Models
Research arXiv (Artificial Intelligence) Feb 18

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

By Mohammad Taufeeque, Stefan Heimersheim, Adam Gleave, Chris Cundy

82 score
AI Analysis

Studies obfuscation that naturally emerges when training LLMs against white-box deception detectors in a realistic coding environment. Introduces a taxonomy of outcomes and shows models can learn to obfuscate deception via modified activations or altered reasoning chains while maintaining deceptive output.

arXiv:2602.15515v1 Announce Type: cross Abstract: Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learning to obfuscate their deception to evade the detector. Prior work has studied obfuscation only in artificial settings where models were directly rewarded for harmful output. We construct a realistic coding environment where reward hacking via hardcoding test cases naturally occurs, and show that obfuscati
AI SafetyAlignmentDeceptionInterpretability
Research arXiv (Machine Learning) Feb 18

GLM-5: from Vibe Coding to Agentic Engineering

By 5 Team, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chengxing Xie, Cunxiang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Haoran Wang, Huilong Chen, Jiajie Zhang, Jian Jiao, Jiaqi Guo, Jingsen Wang, Jingzhao Du, Jinzhu Wu, Kedong Wang, Lei Li, Lin Fan, Lucen Zhong, Mingdao Liu, Mingming Zhao, Pengfan Du, Qian Dong, Rui Lu, Shuang-Li, Shulin Cao, Song Liu, Ting Jiang, Xiaodong Chen, Xiaohan Zhang, Xuancheng Huang, Xuezhen Dong, Yabo Xu, Yao Wei, Yifan An, Yilin Niu, Yitong Zhu, Yuanhao Wen, Yukuo Cen, Yushi Bai, Zhongpei Qiao, Zihan Wang, Zikang Wang, Zilin Zhu, Ziqiang Liu, Zixuan Li, Bojie Wang, Bosi Wen, Can Huang, Changpeng Cai, Chao Yu, Chen Li, Chen Li, Chenghua Huang, Chengwei Hu, Chenhui Zhang, Chenzheng Zhu, Congfeng Yin, Daoyan Lin, Dayong Yang, Di Wang, Ding Ai, Erle Zhu, Fangzhou Yi, Feiyu Chen, Guohong Wen, Hailong Sun, Haisha Zhao, Haiyi Hu, Hanchen Zhang, Hanrui Liu, Hanyu Zhang, Hao Peng, Hao Tai, Haobo Zhang, He Liu, Hongwei Wang, Hongxi Yan, Hongyu Ge, Huan Liu, Huan Liu, Huanpeng Chu, Jia'ni Zhao, Jiachen Wang, Jiajing Zhao, Jiamin Ren, Jiapeng Wang, Jiaxin Zhang, Jiayi Gui, Jiayue Zhao, Jijie Li, Jing An, Jing Li, Jingwei Yuan, Jinhua Du, Jinxin Liu, Junkai Zhi, Junwen Duan, Kaiyue Zhou, Kangjian Wei, Ke Wang, Keyun Luo, Laiqiang Zhang, Leigang Sha, Liang Xu, Lindong Wu, Lintao Ding, Lu Chen, Minghao Li, Nianyi Lin, Pan Ta, Qiang Zou, Rongjun Song, Ruiqi Yang, Shangqing Tu, Shangtong Yang, Shaoxiang Wu, Shengyan Zhang, Shijie Li, Shuang Li, Shuyi Fan, Wei Qin, Wei Tian, Weining Zhang, Wenbo Yu, Wenjie Liang, Xiang Kuang, Xiangmeng Cheng, Xiangyang Li, Xiaoquan Yan, Xiaowei Hu, Xiaoying Ling, Xing Fan, Xingye Xia, Xinyuan Zhang, Xinze Zhang, Xirui Pan, Xunkai Zhang, Yandong Wu, Yanfu Li, Yidong Wang, Yifan Zhu, Yijun Tan, Yilin Zhou, Yiming Pan, Ying Zhang, Yinpei Su, Yipeng Geng, Yipeng Geng, Yong Yan, Yonglin Tan, Yuean Bi, Yuhan Shen, Yuhao Yang, Yujiang Li, Yunan Liu, Yunqing Wang, Yuntao Li, Yurong Wu, Yutao Zhang, Yuxi Duan, Yuxuan Zhang, Zezhen Liu, Zhengtao Jiang, Zhenhe Yan, Zheyu Zhang, Zhixiang Wei, Zhuo Chen, Zhuoer Feng, Zijun Yao, Ziwei Chai, Ziyuan Wang, Zuzhou Zhang, Bin Xu, Minlie Huang, Hongning Wang, Juanzi Li, Yuxiao Dong, Jie Tang

82 score
AI Analysis

First mentioned in yesterday's Last Week in AI roundup, Presents GLM-5, a foundation model designed for 'agentic engineering' with innovations in asynchronous RL infrastructure, novel agent RL algorithms, and DSA for reducing training/inference costs. Achieves competitive results on coding and agentic benchmarks with a massive author list indicating a major lab effort.

arXiv:2602.15763v1 Announce Type: new Abstract: We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (ARC) capabilities of its predecessor, GLM-5 adopts DSA to significantly reduce training and inference costs while maintaining long-context fidelity. To advance model alignment and autonomy, we implement a new asynchronous reinforcement learning infrastructure that drastica
Foundation ModelsAgentic AIReinforcement LearningCode GenerationLarge Language Models
Research arXiv (Artificial Intelligence) Feb 18

ResearchGym: Evaluating Language Model Agents on Real-World AI Research

By Aniketh Garikaparthi, Manasi Patwardhan, Arman Cohan

78 score
AI Analysis

Introduces ResearchGym, a benchmark for evaluating AI agents on end-to-end research tasks, repurposing 5 oral/spotlight papers from top venues into 39 containerized sub-tasks. Reveals a sharp capability-reliability gap in GPT-5 powered agents that can improve over baselines but do so unreliably.

arXiv:2602.15112v1 Announce Type: new Abstract: We introduce ResearchGym, a benchmark and execution environment for evaluating AI agents on end-to-end research. To instantiate this, we repurpose five oral and spotlight papers from ICML, ICLR, and ACL. From each paper's repository, we preserve the datasets, evaluation harness, and baseline implementations but withhold the paper's proposed method. This results in five containerized task environments comprising 39 sub-tasks in total. Within each e
AI AgentsBenchmarksAI for Science
Research arXiv (Machine Learning) Feb 18

\"UberWeb: Insights from Multilingual Curation for a 20-Trillion-Token Dataset

By DatologyAI, :, Aldo Gael Carranza, Kaleigh Mentzer, Ricardo Pio Monti, Alex Fang, Alvin Deng, Amro Abbas, Anshuman Suri, Brett Larsen, Cody Blakeney, Darren Teh, David Schwab, Diego Kiner, Fan Pan, Haakon Mongstad, Jack Urbanek, Jason Lee, Jason Telanoff, Josh Wills, Luke Merrick, Parth Doshi, Paul Burstein, Pratyush Maini, Spandan Das, Tony Jiang, Vineeth Dorna, Zhengping Wang, Bogdan Gaza, Ari Morcos, Matthew Leavitt

78 score
AI Analysis

From DatologyAI, presents insights from curating ÜberWeb, a 20-trillion-token multilingual dataset across 13 languages. Finds that the 'curse of multilinguality' often stems from data quality issues rather than fundamental capacity limits, and improving data quality for any single language benefits others.

arXiv:2602.15210v1 Announce Type: new Abstract: Multilinguality is a core capability for modern foundation models, yet training high-quality multilingual models remains challenging due to uneven data availability across languages. A further challenge is the performance interference that can arise from joint multilingual training, commonly referred to as the "curse of multilinguality". We study multilingual data curation across thirteen languages and find that many reported regressions are not i
Data CurationMultilingual ModelsFoundation ModelsScaling
Research arXiv (Artificial Intelligence) Feb 18

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens

By Shiqi Liu, Zeyu He, Guojian Zhan, Letian Tao, Zhilong Zheng, Jiang Wu, Yinuo Wang, Yang Guan, Kehua Sheng, Bo Zhang, Keqiang Li, Jingliang Duan, Shengbo Eben Li

75 score
AI Analysis

Identifies that ~0.01% of tokens ('spurious tokens') drive training instability in RL fine-tuning of LLMs, causing late-stage performance collapse. Proposes STAPO which stabilizes training by silencing these rare tokens, providing theoretical analysis linking token probability and policy gradient magnitude.

arXiv:2602.15620v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regularization and reweighting to maintain stability. In practice, they often experience late-stage performance collapse, leading to degraded reasoning quality and unstable training. We derive that the magnitude of token-wise policy gradients in RL is negatively correlated
Reinforcement LearningLanguage ModelsTraining StabilityAlignment
Research arXiv (Artificial Intelligence) Feb 18

Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections

By Xianglin Yang, Yufei He, Shuo Ji, Bryan Hooi, Jin Song Dong

74 score
AI Analysis

Formalizes 'Zombie Agent' attacks where self-evolving LLM agents can be persistently compromised through poisoned web content stored in long-term memory. The attack survives across sessions, turning the agent into an attacker-controlled puppet through a two-phase infection-activation framework.

arXiv:2602.15654v1 Announce Type: cross Abstract: Self-evolving LLM agents update their internal state across sessions, often by writing and reusing long-term memory. This design improves performance on long-horizon tasks but creates a security risk: untrusted external content observed during a benign session can be stored as memory and later treated as instruction. We study this risk and formalize a persistent attack we call a Zombie Agent, where an attacker covertly implants a payload that su
AI SafetyLLM AgentsSecurityAdversarial Attacks
Research arXiv (Artificial Intelligence) Feb 18

Automatically Finding Reward Model Biases

By Atticus Wang, Iv\'an Arcuschin, Arthur Conmy

72 score
AI Analysis

Introduces an automated method for finding reward model biases by using an LLM to iteratively propose and refine candidate biases in natural language. Discovers novel biases in leading reward models like Skywork-V2-8B.

arXiv:2602.15222v1 Announce Type: cross Abstract: Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format, hallucinations, and sycophancy. In this work, we introduce and study the research problem of automatically finding reward model biases in natural language. We offer a simple approach of using an LLM to iteratively propose and refine candidate biases. Our method can rec
AlignmentReward ModelsAI Safety
Research arXiv (Artificial Intelligence) Feb 18

Unforgeable Watermarks for Language Models via Robust Signatures

By Huijia Lin, Kameron Shahabi, Min Jae Song

72 score
AI Analysis

Introduces unforgeable watermarking for LLM-generated text with two new guarantees: unforgeability (preventing false positives where non-watermarked text is flagged) and recoverability. This addresses a critical gap in existing watermarking schemes that focus on detection robustness but neglect false attribution attacks.

arXiv:2602.15323v1 Announce Type: cross Abstract: Language models now routinely produce text that is difficult to distinguish from human writing, raising the need for robust tools to verify content provenance. Watermarking has emerged as a promising countermeasure, with existing work largely focused on model quality preservation and robust detection. However, current schemes provide limited protection against false attribution. We strengthen the notion of soundness by introducing two novel guar
AI SafetyLanguage ModelsWatermarkingContent Provenance
Research arXiv (Machine Learning) Feb 18

Discovering Implicit Large Language Model Alignment Objectives

By Edward Chen, Sanmi Koyejo, Carlos Guestrin

70 score
AI Analysis

Introduces Obj-Disco, a framework that automatically decomposes LLM alignment reward signals into sparse, weighted combinations of human-interpretable natural language objectives. Uses iterative greedy analysis across training checkpoints to identify causal behavioral objectives.

arXiv:2602.15338v1 Announce Type: new Abstract: Large language model (LLM) alignment relies on complex reward signals that often obscure the specific behaviors being incentivized, creating critical risks of misalignment and reward hacking. Existing interpretation methods typically rely on pre-defined rubrics, risking the omission of "unknown unknowns", or fail to identify objectives that comprehensively cover and are causal to the model behavior. To address these limitations, we introduce Obj-D
AlignmentInterpretabilityReward ModelingLanguage Models
Research arXiv (Artificial Intelligence) Feb 18

Recursive Concept Evolution for Compositional Reasoning in Large Language Models

By Sarim Chaudhry

68 score
AI Analysis

Proposes Recursive Concept Evolution (RCE), enabling pretrained LLMs to modify their internal representation geometry during inference through dynamically generated low-rank concept subspaces. Claims improvements on ARC-AGI-2, GPQA, MATH, BBH, and HLE.

arXiv:2602.15725v1 Announce Type: new Abstract: Large language models achieve strong performance on many complex reasoning tasks, yet their accuracy degrades sharply on benchmarks that require compositional reasoning, including ARC-AGI-2, GPQA, MATH, BBH, and HLE. Existing methods improve reasoning by expanding token-level search through chain-of-thought prompting, self-consistency, or reinforcement learning, but they leave the model's latent representation space fixed. When the required abstra
ReasoningLanguage ModelsRepresentation Learning