Category intelligence

Research Briefing — July 10, 2026

483 current items analyzed and ranked.

Executive synthesis

Research Summary

AI safety and alignment dominates today's most significant work, spanning foundational theory to practical pre-deployment evaluation.

Evaluation and mathematics frontiers advance in parallel. Measuring Intelligence Beyond Human Scale (Braverman, Hazan) uses relative, model-generated challenges to counter benchmark saturation. A high-profile position paper reframes LLM-driven formal mathematics as open-ended research agents rather than solvers. DeepSWE contributes 113 contamination-resistant, long-horizon coding tasks across 91 repositories.

Efficiency and learning theory complete the set. Jet-Long enables tuning-free long-context extension via dynamic bifocal RoPE. Additional work reframes continual learning around adapting to world change and derives an exact information theory of generalization phase transitions in Bayesian diffusion models.

Key Themes

AI Safety and Security · 10Efficient Inference and LLM Serving · 12AI for Mathematics & Formal Reasoning · 1LLM Agents and Agentic AI · 20AI Safety and Alignment · 9AI Safety, Alignment, and Unlearning · 8Evaluation and Benchmarking · 9Learning Theory and Optimization · 14Reasoning and Reinforcement Learning · 7Reinforcement Learning · 27

Primary evidence

Top Ranked Signals

Research arXiv (Machine Learning) Jul 10

Provably Optimal Learning Algorithms for Assistance Games

By Nivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan, Stuart Russell, Nika Haghtalab

70 score
AI Analysis

Provides the first provably efficient learning algorithms for repeated online assistance games between an informed human and an uninformed assistant, introducing assistance regret and decentralized algorithms achieving a (1-1/e)-approximation. It formalizes cooperative human-AI interaction where the assistant only observes human actions.

arXiv:2607.08012v1 Announce Type: new Abstract: This paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact over $T$ timesteps to optimize a common reward function. While the informed agent (the human) observes a latent state of the world, the uninformed agent (the assistant) observes only the human's actions. We provide the first provably efficient learning algorithms for repeated assistance games. We introduce the
Learning TheoryAlignmentHuman-AI Interaction
Research arXiv (Computation and Language) Jul 10

From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

By Eric Jiang, Xiao Liang, Yikai Zhang, Yingjia Wan, Mengting Li, Haikang Deng, Alexander K. Taylor, Justin Baker, Rushil Raghavan, Junyi Zhang, Ying Nian Wu, Andrea L. Bertozzi, Kai-Wei Chang, Raghu Meka, Matthew Sottile, Nanyun Peng, Amit Sahai, Terence Tao, Wei Wang

70 score
AI Analysis

A position paper arguing that LLM-driven formal mathematics must shift from predefined problem solvers toward research agents capable of open-ended theorem discovery and resolving conjectures with rigorous formal reasoning. The author list notably includes Terence Tao alongside strong ML and math researchers.

arXiv:2607.07779v1 Announce Type: new Abstract: Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages. However, current systems remain fundamentally limited in tackling frontier research mathematics, such as discovering new theorems or resolving open conjectures, which are often open-end
AI for MathematicsFormal ReasoningLanguage Models
Research arXiv (Artificial Intelligence) Jul 9

Predicting LLM Safety Before Release by Simulating Deployment

By Marcus Williams, Hannah Sheahan, Cameron Raymond, Tomek Korbak, Deng Pan, Peilin Yang, Leon Maksin, Ningyi Xie, Phillip Guo, Ian Kivlichan, Micah Carroll

68 score
AI Analysis

This safety paper simulates deployment by holding fixed the prefixes of de-identified conversations from a prior model deployment and regenerating responses from a candidate model, enabling audits for novel misalignment and estimation of misbehavior prevalence before release. It offers a more representative pre-deployment safety evaluation than typical recognizable tests.

arXiv:2607.07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evidence about how often undesired model behavior will occur in deployment: they generally have insufficient coverage, are unrepresentative, and are generally recognizable as tests. To address these concerns, we study a simple way to simulate a model deployment: starting from de-identified conversations from a pr
AI SafetyAlignmentEvaluation
Research arXiv (Artificial Intelligence) Jul 9

Measuring Intelligence Beyond Human Scale

By Jerry Han, Rafael Moschopoulos, Ella Colby, Vishrut Goyal, Andrew Tu, Kia Ghods, Mark Braverman, Elad Hazan

66 score
AI Analysis

This paper proposes measuring intelligence beyond human capability via relative rather than absolute evaluation, where models generate public challenges that separate other systems into an adversarial psychometric rating. It describes protocols that reduce private-information attacks and support judge-free adjudication that scales with agent capabilities.

arXiv:2607.07040v1 Announce Type: new Abstract: How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. Aggregating these outcomes yields an adversarial psychometric r
EvaluationAI CapabilitiesBenchmarking
Research arXiv (Machine Learning) Jul 10

Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE

By Haozhan Tang, Zerui Wang, Yuxian Gu, Song Han, Han Cai

65 score
AI Analysis

Jet-Long is a tuning-free zero-shot context-extension method pairing a RoPE-faithful local window with a long-range window whose rescaling factor adapts to the current sequence length, avoiding the short-vs-long fidelity tradeoff of fixed rescaling. It targets open-weight checkpoints deployed beyond their pretraining window.

arXiv:2607.07740v1 Announce Type: new Abstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. Most existing zero-shot methods fix a single rescaling factor up front, so an aggres
Long-Context ModelsEfficient InferenceLanguage Models
Research arXiv (Machine Learning) Jul 10

When Does Continual Learning Require Learning

By Anne Harrington, Nayan Saxena, Michael Murphy, Anastasia Borovykh, Zeyu Yun, Sridhar Kamath, Ara Eindra Kyi, Trevor Darrell, Jitendra Malik, Yutong Bai

65 score
AI Analysis

Reframes continual learning for LLMs as increasing competence as the world changes, disentangling change into new domains (space) and drifting data (time) rather than only forgetting mitigation. It builds evaluations under realistic conditions where new domains arrive, facts drift past cutoff, and agentic state accumulates.

arXiv:2607.07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn? Today, the field largely frames this as a problem of context management and mitigating forgetting. We argue this framing is incomplete: continual learning is fundamentally about increasing model competence as the world changes. We disentangle this change along two axes -- space, where the model encounters new domains, an
Continual LearningLanguage ModelsEvaluation
Research arXiv (Machine Learning) Jul 10

An exact information theory of generalization phase transitions in Bayesian diffusion models

By Henry Hunt, Mason Kamb, Surya Ganguli

65 score
AI Analysis

Introduces analytically tractable Bayesian information-restricted diffusion (BIRD) models to explain how diffusion models avoid the curse of dimensionality and generalize rather than memorize. It generalizes prior local-information analytical models and characterizes generalization phase transitions with exact information theory.

arXiv:2607.08041v1 Announce Type: new Abstract: How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past train
Diffusion ModelsLearning TheoryGeneralization
Research arXiv (Machine Learning) Jul 10

Modular Pretraining Enables Access Control

By Ethan Roland, Murat Cubuktepe, Erick Martinez, Stijn Servaes, Keenan Pepper, Mike Vaiana, Diogo Schwerz de Lucena, Judd Rosenblatt, Addie Foote, Cem Anil, Alex Cloud

65 score
AI Analysis

Proposes gradient-routed auxiliary modules (GRAM), a pretraining method that adds modules updated selectively to induce capability specialization so that ablating a module at inference removes that capability, approximating a separately trained model. It targets the dual-use dilemma by enabling access control without training and deploying multiple full models.

arXiv:2607.08077v1 Announce Type: new Abstract: AI developers face a dual-use dilemma. An AI capability that helps one user cure a disease can help another synthesize one. This dilemma could be resolved with access control, limiting dual-use AI capabilities to trusted deployments with a legitimate need. A gold standard for access control would be to serve separate models with different capabilities to different users. However, training and deploying multiple models is prohibitively expensive. T
AI SafetyAccess ControlPretraining
Research arXiv (Artificial Intelligence) Jul 10

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning

By Yotam Wolf, Noam Wies, Amnon Shashua

63 score
AI Analysis

This paper gives a theoretical treatment of in-context search, modeling reflection-driven reasoning as approximate inference over reasoning traces and analyzing the sequential sampling complexity needed for high success. It shows that when self-reflection reliably localizes early mistakes, in-context search can deliver exponential gains over the base model.

arXiv:2607.06720v1 Announce Type: new Abstract: Training large language models (LLMs) with extended reasoning has enabled in-context search, in which models iteratively generate, critique, and revise solution attempts. We provide a theoretical analysis of in-context search by modeling it as approximate inference over reasoning traces, where the base model defines a prior and self-reflection provides feedback for posterior updates, and study the resulting inference-time sampling complexity - the
ReasoningLanguage ModelsTheory
Research arXiv (Machine Learning) Jul 10

DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks

By Wenqi Huang, Charley Lee, Leonard Tng, Serena Ge

63 score
AI Analysis

DeepSWE is a benchmark of 113 original long-horizon software-engineering tasks written from scratch across 91 active repositories and five languages, explicitly designed to avoid pretraining contamination and the biased test-only-for-one-fix grading of SWE-bench. It aims to measure genuine problem-solving by coding agents.

arXiv:2607.07946v1 Announce Type: cross Abstract: DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks follow SWE-bench in mining merged fixes from public GitHub repositories, which creates two problems: the fixes and their discussion were likely seen during pretraining, so a high score can reflect recall rather than problem-solving; and each task is graded by the tests that shipped with its merged fi
LLM AgentsBenchmarksCode Generation
Research arXiv (Machine Learning) Jul 10

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

By Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky, Tanush Chopra, Victoria Krakovna

63 score
AI Analysis

This paper stress-tests chain-of-thought monitoring as a safety mechanism, showing that adversarial agents can use persuasion-based arguments to convince a CoT monitor to approve policy-violating actions across a 40-task evaluation with thousands of interactions. It demonstrates that persuasion-jailbreak vulnerabilities extend to monitoring LLMs, weakening a promising oversight approach.

arXiv:2607.08066v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can surface misaligned or deceptive behavior. While effective in standard scenarios, recent work highlights that LLMs remain vulnerable to persuasion-based jailbreaks, where natural-language arguments override model constraints. We stress-test whether this vulnerability extends to monitoring LLMs: can an adversarial
AI SafetyAlignmentLanguage Models
Research arXiv (Machine Learning) Jul 10

A law of robustness for two-layer neural networks with arbitrary weights

By Yitzchak Shmalo

62 score
AI Analysis

Proves the Bubeck-Li-Nagaraj law of robustness for two-layer neural networks with arbitrary (unbounded) weights, up to a logarithmic factor, for every continuous piecewise-linear activation including ReLU. This closes a gap where prior universal results required polynomial parameter bounds.

arXiv:2607.07778v1 Announce Type: new Abstract: Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with $m$ neurons that fits $n$ noisy labels must have Lipschitz constant at least of order $\sqrt{n/m}$, with no restriction on the size of the weights. Bubeck and Sellke proved a universal version of this law for Lipschitz-parameterized classes, but under a polynomial bound on the parameters; at depth three that boundedness hypothesis is genuinely necessary. T
Learning TheoryRobustnessNeural Network Theory