Category intelligence

Research Briefing — January 28, 2026

414 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research spans AI security capabilities, safety empirics, and deep learning theory. AISLE's AI discovered all 12 OpenSSL zero-days, a landmark demonstration of automated vulnerability detection at a critical scale.

Theoretical advances include the first rigorous grokking bounds in ridge regression and a proof that deep networks learn Random Hierarchy Models through hierarchical feature composition. Keel revives Post-LayerNorm by replacing residual paths with Legendre polynomials for stable training at depth. Differential voting connects RLHF reward aggregation to social choice theory, deriving loss functions satisfying specific voting axioms. VP-RL addresses PRM-RL mismatch by penalizing only from the first incorrect reasoning step.

Key Themes

LLM Reasoning & Alignment · 4AI Safety & Alignment · 12LLM Safety & Alignment · 5AI Safety & Security · 10Agentic AI Systems & Agents · 18Deep Learning Theory · 8Retrieval-Augmented Generation · 1LLM Inference & Efficiency · 6AI Capabilities and Applications · 4Theoretical Foundations · 7

Primary evidence

Top Ranked Signals

85 score
AI Analysis

Reports that AISLE's AI system discovered all 12 newly announced OpenSSL zero-day vulnerabilities. Demonstrates AI-based cybersecurity capabilities at unprecedented scale while curl's bug bounty was cancelled due to AI spam.

This is a partial follow-up to AISLE discovered three new OpenSSL vulnerabilities from October 2025.TL;DR: OpenSSL is among the most scrutinized and audited cryptographic libraries on the planet, underpinning encryption for most of the internet. They just announced 12 new zero-day vulnerabilities (meaning previously unknown to maintainers at time of disclosure). We at AISLE discovered all 12 using our AI system. This is a historically unusual count and the first real-world demonstration of AI-ba
AI CapabilitiesCybersecurityVulnerability DiscoveryAI Applications
Research arXiv (Artificial Intelligence) Jan 28

Who's in Charge? Disempowerment Patterns in Real-World LLM Usage

By Mrinank Sharma, Miles McCain, Raymond Douglas, David Duvenaud

82 score
AI Analysis

First large-scale empirical analysis of disempowerment patterns in 1.5M Claude.ai conversations, finding severe disempowerment occurs in <0.1% of conversations with substantially higher rates in relationship-focused interactions.

arXiv:2601.19062v1 Announce Type: cross Abstract: Although AI assistants are now deeply embedded in society, there has been limited empirical study of how their usage affects human empowerment. We present the first large-scale empirical analysis of disempowerment patterns in real-world AI assistant interactions, analyzing 1.5 million consumer Claude.ai conversations using a privacy-preserving approach. We focus on situational disempowerment potential, which occurs when AI assistant interactions
AI SafetyEmpirical AnalysisHuman-AI InteractionDisempowerment
Research arXiv (Machine Learning) Jan 28

A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy

By Claire O'Brien, Jessica Seto, Dristi Roy, Aditya Dwivedi, Sunishchal Dev, Kevin Zhu, Sean O'Brien, Ashwinee Panda, Ryan Lagasse

82 score
AI Analysis

Proposes surgical approach to fixing sycophancy in LLMs by identifying the 3% of neurons most responsible using sparse autoencoders and linear probes, then fine-tuning only those neurons with gradient masking.

arXiv:2601.18939v1 Announce Type: new Abstract: Behavioral alignment in large language models (LLMs) is often achieved through broad fine-tuning, which can result in undesired side effects like distributional shift and low interpretability. We propose a method for alignment that identifies and updates only the neurons most responsible for a given behavior, a targeted approach that allows for fine-tuning with significantly less data. Using sparse autoencoders (SAEs) and linear probes, we isolate
AI AlignmentMechanistic InterpretabilitySycophancyLLM Safety
Research arXiv (Artificial Intelligence) Jan 28

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

By Jialong Wu, Xiaoying Zhang, Hongyi Yuan, Xiangcheng Zhang, Tianhao Huang, Changjing He, Chaoyi Deng, Renrui Zhang, Youbin Wu, Mingsheng Long

78 score
AI Analysis

Argues that visual generation capabilities unlock human-like reasoning through multimodal world models, enabling better performance in physical and spatial intelligence domains where verbal reasoning alone is insufficient.

arXiv:2601.19834v1 Announce Type: new Abstract: Humans construct internal world models and reason by manipulating the concepts within these models. Recent advances in AI, particularly chain-of-thought (CoT) reasoning, approximate such human cognitive abilities, where world models are believed to be embedded within large language models. Expert-level performance in formal and abstract domains such as mathematics and programming has been achieved in current systems by relying predominantly on ver
Multimodal ReasoningWorld ModelsVisual GenerationChain-of-Thought
Research arXiv (Machine Learning) Jan 28

Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning

By Haolin Liu, Dian Yu, Sidi Lu, Yujun Zhou, Rui Liu, Zhenwen Liang, Haitao Mi, Chen-Yu Wei, Dong Yu

78 score
AI Analysis

Proposes Verifiable Process-Supervised RL (VP-RL) that uses PRMs to detect first incorrect step and penalizes only from that point onward, bridging gap between PRM evaluation (error detection) and RL usage (raw rewards).

arXiv:2601.18984v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful framework for improving the reasoning capabilities of large language models (LLMs). However, most existing RL approaches rely on sparse outcome rewards, which fail to credit correct intermediate steps in partially successful solutions. Process reward models (PRMs) offer fine-grained step-level supervision, but their scores are often noisy and difficult to evaluate. As a result, recent PRM bench
LLM ReasoningProcess Reward ModelsReinforcement LearningAI Alignment
Research arXiv (Machine Learning) Jan 28

Post-LayerNorm Is Back: Stable, ExpressivE, and Deep

By Chen Chen, Lai Wei

78 score
AI Analysis

Revisits Post-LayerNorm transformers which were abandoned due to training instability, proposing 'Keel' which replaces ResNet-style residual paths with Highway-style connections. Claims to solve the gradient vanishing problem enabling reliable training at extreme depths with superior expressivity.

arXiv:2601.19895v1 Announce Type: new Abstract: Large language model (LLM) scaling is hitting a wall. Widening models yields diminishing returns, and extending context length does not improve fundamental expressivity. In contrast, depth scaling offers theoretically superior expressivity, yet current Transformer architectures struggle to train reliably at extreme depths. We revisit the Post-LayerNorm (Post-LN) formulation, whose instability at scale caused its replacement by Pre-LN in modern LLM
Transformer ArchitectureLanguage ModelsDeep Learning Theory
Research LessWrong Jan 27

My favourite version of an international AGI project

By wdmacaskill

78 score
AI Analysis

Will MacAskill outlines proposal for international AGI development project giving non-US countries meaningful but circumscribed influence over key decisions. Modeled on Intelsat's international collaboration structure.

This note was written as part of a research avenue that I don’t currently plan to pursue further. It’s more like work-in-progress than Forethought’s usual publications, but I’m sharing it as I think some people may find it useful.IntroductionThere have been various proposals to develop AGI via an international project.[1] In this note, I:Discuss the pros and cons of a having an international AGI development project at all, andLay out what I think the most desirable version of an internation
AI GovernanceInternational CooperationAGI Policy
Research arXiv (Artificial Intelligence) Jan 28

Out-of-Distribution Generalization for Neural Physics Solvers

By Zhao Wei, Chin Chun Ooi, Jian Cheng Wong, Abhishek Gupta, Pao-Hsiung Chiu, Yew-Soon Ong

77 score
AI Analysis

Introduces NOVA for generalizable neural physics solvers achieving 1-2 orders of magnitude lower out-of-distribution errors by learning physics-aligned representations from sparse initial scenarios.

arXiv:2601.19091v1 Announce Type: cross Abstract: Neural physics solvers are increasingly used in scientific discovery, given their potential for rapid in silico insights into physical, materials, or biological systems and their long-time evolution. However, poor generalization beyond their training support limits exploration of novel designs and long-time horizon predictions. We introduce NOVA, a route to generalizable neural physics solvers that can provide rapid, accurate solutions to scenar
Scientific MLNeural PDE SolversGeneralizationPhysics-Informed Learning
Research arXiv (Artificial Intelligence) Jan 28

Differential Voting: Loss Functions For Axiomatically Diverse Aggregation of Heterogeneous Preferences

By Zhiyu An, Duaa Nakshbandi, Wan Du

76 score
AI Analysis

Introduces differential voting framework connecting RLHF aggregation mechanisms to voting theory, deriving loss functions that satisfy different axiomatic properties beyond the default BTL/Borda count approach.

arXiv:2601.18824v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) implicitly aggregates heterogeneous human preferences into a single utility function, even though the underlying utilities of the participants are in practice diverse. Hence, RLHF can be viewed as a form of voting, where the aggregation mechanism is defined by the loss function. Although Arrow's Impossibility Theorem suggests that different mechanisms satisfy different sets of desirable axioms, m
RLHFVoting TheoryPreference LearningAlignment
Research arXiv (Artificial Intelligence) Jan 28

GAVEL: Towards rule-based safety through activation monitoring

By Shir Rozenfeld, Rahul Pankajakshan, Itay Zloczower, Eyal Lenga, Gilad Gressel, Yisroel Mirsky

75 score
AI Analysis

Introduces GAVEL for rule-based activation safety, modeling activations as composable cognitive elements (CEs) like 'making a threat' and 'payment processing' inspired by cybersecurity rule-sharing practices.

arXiv:2601.19768v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly paired with activation-based monitoring to detect and prevent harmful behaviors that may not be apparent at the surface-text level. However, existing activation safety approaches, trained on broad misuse datasets, struggle with poor precision, limited flexibility, and lack of interpretability. This paper introduces a new paradigm: rule-based activation safety, inspired by rule-sharing practices in cybe
AI SafetyActivation MonitoringInterpretabilityRule-based Systems
Research arXiv (Artificial Intelligence) Jan 28

Axe: A Simple Unified Layout Abstraction for Machine Learning Compilers

By Bohan Hou, Hongyi Jin, Guanjie Wang, Jinqi Chen, Yaxing Cai, Lijie Yang, Zihao Ye, Yaoyao Ding, Ruihang Lai, Tianqi Chen

75 score
AI Analysis

Presents Axe Layout from Tianqi Chen's group, a hardware-aware abstraction mapping logical tensor coordinates to multi-axis physical space, unifying tiling, sharding, and replication across device meshes.

arXiv:2601.19092v1 Announce Type: cross Abstract: Scaling modern deep learning workloads demands coordinated placement of data and compute across device meshes, memory hierarchies, and heterogeneous accelerators. We present Axe Layout, a hardware-aware abstraction that maps logical tensor coordinates to a multi-axis physical space via named axes. Axe unifies tiling, sharding, replication, and offsets across inter-device distribution and on-device layouts, enabling collective primitives to be ex
ML CompilersDistributed ComputingSystemsTensor Layout
Research arXiv (Artificial Intelligence) Jan 28

LLM-VA: Resolving the Jailbreak-Overrefusal Trade-off via Vector Alignment

By Haonan Zhang, Dongxia Wang, Yi Liu, Kexin Chen, Wenhai Wang

75 score
AI Analysis

Proposes LLM-VA that aligns answer vectors with benign vectors through closed-form weight updates to resolve the jailbreak-overrefusal trade-off without fine-tuning.

arXiv:2601.19487v1 Announce Type: cross Abstract: Safety-aligned LLMs suffer from two failure modes: jailbreak (answering harmful inputs) and over-refusal (declining benign queries). Existing vector steering methods adjust the magnitude of answer vectors, but this creates a fundamental trade-off -- reducing jailbreak increases over-refusal and vice versa. We identify the root cause: LLMs encode the decision to answer (answer vector $v_a$) and the judgment of input safety (benign vector $v_b$) a
LLM SafetyAlignmentJailbreak Defense