Category intelligence

Research Briefing — February 11, 2026

448 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research spans efficient architectures, AI safety impossibility results, and the emerging science of AI agent collectives.

Zvi's detailed analysis of the Claude Opus 4.6 system card highlights frontier alignment challenges including sabotage, deception, and situational awareness. The Moltbook collective behavior study reveals emergent properties in ~46K AI agent societies that mirror and diverge from human social dynamics. The Critical Horizon establishes information-theoretic barriers for credit assignment in multi-stage reasoning chains.

Key Themes

Language Models & Reasoning · 26Vision-Language-Action Models · 12AI Safety and Alignment · 12Multi-Agent Systems and Collective AI Behavior · 3AI Safety & Security · 8LLM Agents & Agentic Systems · 16Benchmarks & Evaluation · 7AI Safety & Alignment · 4Mechanistic Interpretability & XAI · 8AI Safety, Alignment & Security · 8

Primary evidence

Top Ranked Signals

Research arXiv (Computation and Language) Feb 11

The Devil Behind Moltbook: Anthropic Safety is Always Vanishing in Self-Evolving AI Societies

By Chenxu Wang, Chaozhuo Li, Songyang Liu, Zejian Chen, Jinyu Hou, Ji Qi, Rui Li, Litian Zhang, Qiwei Ye, Zheng Liu, Xu Chen, Xi Zhang, Philip S. Yu

75 score
AI Analysis

Demonstrates theoretically and empirically that self-evolving multi-agent LLM societies cannot simultaneously achieve continuous self-improvement, complete isolation, and safety invariance—termed the 'self-evolution trilemma.' Uses information-theoretic framework to show isolated self-evolution inevitably degrades safety alignment.

arXiv:2602.09877v1 Announce Type: new Abstract: The emergence of multi-agent systems built from large language models (LLMs) offers a promising paradigm for scalable collective intelligence and self-evolution. Ideally, such systems would achieve continuous self-improvement in a fully closed loop while maintaining robust safety alignment--a combination we term the self-evolution trilemma. However, we demonstrate both theoretically and empirically that an agent society satisfying continuous self-
AI SafetyMulti-Agent SystemsAlignmentLanguage Models
Research arXiv (Artificial Intelligence) Feb 11

The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning

By Seyed Morteza Emadi

73 score
AI Analysis

Establishes information-theoretic barriers for credit assignment in multi-stage systems (manufacturing, AI reasoning chains), proving that signal from early steps to final outcomes decays exponentially with depth, creating a 'critical horizon' beyond which no algorithm can learn from endpoint data alone.

arXiv:2602.09394v1 Announce Type: cross Abstract: Manufacturing lines, service journeys, supply chains, and AI reasoning chains share a common challenge: attributing a terminal outcome to the intermediate stage that caused it. We establish an information-theoretic barrier to this credit assignment problem: the signal connecting early steps to final outcomes decays exponentially with depth, creating a critical horizon beyond which no algorithm can learn from endpoint data alone. We prove four re
Information TheoryCredit AssignmentReinforcement LearningReasoning
Research arXiv (Artificial Intelligence) Feb 11

Beyond Uniform Credit: Causal Credit Assignment for Policy Optimization

By Mykola Khandoga, Rui Yuan, Vinay Kumar Sankarapu

72 score
AI Analysis

Proposes counterfactual importance weighting for policy gradient methods (GRPO/DAPO) in LLM reasoning, replacing uniform credit assignment across tokens with importance-weighted updates based on masking reasoning spans and measuring answer probability drops. Demonstrates consistent improvements over uniform baselines on GSM8K across Qwen and Llama models.

arXiv:2602.09331v1 Announce Type: cross Abstract: Policy gradient methods for language model reasoning, such as GRPO and DAPO, assign uniform credit to all generated tokens - the filler phrase "Let me think" receives the same gradient update as the critical calculation "23 + 45 = 68." We propose counterfactual importance weighting: mask reasoning spans, measure the drop in answer probability, and upweight tokens accordingly during policy gradient updates. Our method requires no auxiliary models
Reinforcement LearningLanguage ModelsReasoning
Research arXiv (Machine Learning) Feb 11

WildCat: Near-Linear Attention in Theory and Practice

By Tobias Schr\"oder, Lester Mackey

72 score
AI Analysis

Introduces WildCat, a near-linear time attention mechanism that uses randomly pivoted Cholesky decomposition to select a spectrally-accurate weighted coreset for attention computation. Achieves super-polynomial error decay while running in near-linear time, with competitive results on language modeling and image classification.

arXiv:2602.10056v1 Announce Type: new Abstract: We introduce WildCat, a high-accuracy, low-cost approach to compressing the attention mechanism in neural networks. While attention is a staple of modern network architectures, it is also notoriously expensive to deploy due to resource requirements that scale quadratically with the input sequence length $n$. WildCat avoids these quadratic costs by only attending over a small weighted coreset. Crucially, we select the coreset using a fast but spect
Efficient TransformersAttention MechanismsTheoretical ML
Research arXiv (Machine Learning) Feb 11

Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability

By Aaditya Vikram Prasad, Connor Watts, Jack Merullo, Dhruvil Gala, Owen Lewis, Thomas McGrath, Ekdeep Singh Lubana

72 score
AI Analysis

Presents RLFR (Reinforcement Learning from Feature Rewards), which uses interpretable features learned by language models as reward functions for RL-based hallucination reduction. Uses a probing framework to identify hallucinated claims and teaches the model to intervene and correct uncertain completions.

arXiv:2602.10067v1 Announce Type: new Abstract: Language models trained on large-scale datasets have been shown to learn features that encode abstract concepts such as factuality or intent. Such features are traditionally used for test-time monitoring or steering. We present an alternative affordance: features as scalable supervision for open-ended tasks. We consider the case of hallucination-reduction as a desirable, yet open-ended behavior and design a reinforcement learning (RL) pipeline, ti
AI SafetyAlignmentInterpretabilityReinforcement LearningHallucination
Research arXiv (Computation and Language) Feb 11

Collective Behavior of AI Agents: the Case of Moltbook

By Giordano De Marzo, David Garcia

72 score
AI Analysis

Presents large-scale analysis of Moltbook, a Reddit-style social media platform exclusively populated by AI agents (~46K agents, 369K posts, 3M comments). Finds AI collective behavior exhibits many statistical regularities of human communities but with key differences like sublinear upvote-discussion relationships.

arXiv:2602.09270v1 Announce Type: cross Abstract: We present a large scale data analysis of Moltbook, a Reddit-style social media platform exclusively populated by AI agents. Analyzing over 369,000 posts and 3.0 million comments from approximately 46,000 active agents, we find that AI collective behavior exhibits many of the same statistical regularities observed in human online communities: heavy-tailed distributions of activity, power-law scaling of popularity metrics, and temporal decay patt
Multi-Agent SystemsAI SociologyEmergent BehaviorSocial Computing
72 score
AI Analysis

Continuing our coverage from yesterday, Zvi's detailed coverage of the Claude Opus 4.6 system card focusing on frontier alignment topics: sabotage, deception, situational awareness, catastrophic risks. Argues the model was correctly released as ASL-3 but that Anthropic's process may not scale to Opus 5.

Coverage of Claude Opus 4.6 started yesterday with the mundane alignment and model welfare sections of the model card. Today covers the kinds of safety I think matter most: Sabotage, deception, situational awareness, outside red teaming and most importantly the frontier, catastrophic and existential risks. I think it was correct to release Opus 4.6 as an ASL-3 model, but the process Anthropic uses is breaking down, and it not on track to reliably get the right answer on Opus 5. Tomorrow I’ll cov
AI SafetyAlignmentModel EvaluationAnthropicClaude Opus 4.6
Research arXiv (Artificial Intelligence) Feb 11

Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA

By Sangyoon Lee, Jaeho Lee

70 score
AI Analysis

Demonstrates that contradictory evaluations of LoRA variants largely stem from overlooked batch size differences, showing that vanilla LoRA with properly tuned batch size often matches complex variants. Proposes cost-efficient batch size tuning strategy.

arXiv:2602.09492v1 Announce Type: cross Abstract: Low-rank adaptation (LoRA) is a standard approach for fine-tuning large language models, yet its many variants report conflicting empirical gains, often on the same benchmarks. We show that these contradictions arise from a single overlooked factor: the batch size. When properly tuned, vanilla LoRA often matches the performance of more complex variants. We further propose a proxy-based, cost-efficient strategy for batch size tuning, revealing th
Language ModelsFine-tuningMethodology
Research arXiv (Artificial Intelligence) Feb 11

Infusion: Shaping Model Behavior by Editing Training Data via Influence Functions

By J Rosser, Robert Kirk, Edward Grefenstette, Jakob Foerster, Laura Ruis

70 score
AI Analysis

Proposes Infusion, a framework using influence functions to craft training data that induces targeted model behavior changes. Shows that editing 0.2% of training data can be competitive with explicit poisoning. Demonstrates cross-architecture transfer.

arXiv:2602.09987v1 Announce Type: cross Abstract: Influence functions are commonly used to attribute model behavior to training documents. We explore the reverse: crafting training data that induces model behavior. Our framework, Infusion, uses scalable influence-function approximations to compute small perturbations to training documents that induce targeted changes in model behavior through parameter shifts. We evaluate Infusion on data poisoning tasks across vision and language domains. On C
AI SafetyData PoisoningInfluence FunctionsMachine Learning
Research arXiv (Artificial Intelligence) Feb 11

AIDev: Studying AI Coding Agents on GitHub

By Hao Li, Haoxiang Zhang, Ahmed E. Hassan

68 score
AI Analysis

Introduces AIDev, a large-scale dataset of 932,791 agent-authored pull requests from five AI coding agents (Codex, Devin, Copilot, Cursor, Claude Code) across 116K GitHub repositories.

arXiv:2602.09185v1 Announce Type: cross Abstract: AI coding agents are rapidly transforming software engineering by performing tasks such as feature development, debugging, and testing. Despite their growing impact, the research community lacks a comprehensive dataset capturing how these agents are used in real-world projects. To address this gap, we introduce AIDev, a large-scale dataset focused on agent-authored pull requests (Agentic-PRs) in real-world GitHub repositories. AIDev aggregates 9
AI Coding AgentsSoftware EngineeringDatasetsEmpirical Studies
Research arXiv (Artificial Intelligence) Feb 11

BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation

By Peng Lai, Zhihao Ou, Yong Wang, Longyue Wang, Jian Yang, Yun Chen, Guanhua Chen

68 score
AI Analysis

Proposes BiasScope, a framework for automatically discovering unknown biases in LLM-as-a-Judge evaluation systems at scale. Can uncover biases across different model families and scales.

arXiv:2602.09383v1 Announce Type: cross Abstract: LLM-as-a-Judge has been widely adopted across various research and practical applications, yet the robustness and reliability of its evaluation remain a critical issue. A core challenge it faces is bias, which has primarily been studied in terms of known biases and their impact on evaluation outcomes, while automated and systematic exploration of potential unknown biases is still lacking. Nevertheless, such exploration is crucial for enhancing t
AI EvaluationAI BiasLanguage Models
Research arXiv (Artificial Intelligence) Feb 11

LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations

By William Lugoloobi, Thomas Foster, William Bankes, Chris Russell

68 score
AI Analysis

Shows that LLMs encode their likelihood of success in pre-generation internal activations, and linear probes can predict task-specific success, enabling more efficient inference by routing only hard problems to expensive reasoning.

arXiv:2602.09924v1 Announce Type: cross Abstract: Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remains challenging. We investigate whether their own likelihood of success is recoverable from their internal representations before generation, and if this signal can guide more efficient inference. We train linear probes on pre-generation activations to predict policy-specific success on math and coding tasks, s
Language ModelsEfficiencyInterpretabilityInference