Category intelligence

Research Briefing — June 24, 2026

451 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by safety/alignment and agentic AI, alongside a standout clinical deployment and a foundational theory contribution.

Clinical & applied impact: RaDaR, an open-source 32B reasoning LLM for rare disease diagnosis, anchors the day with a randomized AI-physician-assistance trial validating real-world deployability.

Safety & alignment (the strongest theme):

Agents, interpretability & foundations:

Theory & training dynamics: The Geometry Behind Diffusion and Flow Matching elegantly unifies both under Wasserstein-space gradient flows and geodesics, while a plasticity-loss study probes whether scale rescues continual learning in 5M–314M parameter transformers.

Key Themes

AI Safety and Alignment · 11Agentic AI · 16Reasoning and LLM Evaluation · 12Reinforcement Learning · 19AI Safety & Alignment · 11Healthcare and Biomedical AI · 8Interpretability and Causality · 5Generative Models and Diffusion/Flow · 7LLM Efficiency and Systems · 8AI Safety and Security · 6

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) Jun 24

A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial

By Haichao Chen, Songchi Zhou, Zhengyun Zhao, Shikai Hu, Xianghong Jin, Hongwei Ji, Li He, Shuli Li, Yiming Qin, Xin Tan, Runfeng Shi, Yih Chung Tham, Jiaye Zhu, Ye Li, Ye Jin, Longhao Cao, Dawei Li, Honghan Wu, Hongqiu Gu, Guanqiao Li, Tudor Groza, Chunying Li, Dian Zeng, Weihong Yu, Gareth Baynam, Saumya Shekhar Jamuar, Min Shen, Shuyang Zhang, Bin Sheng, Sheng Yu, Tien Yin Wong

76 score
AI Analysis

RaDaR is an open-source compact 32B reasoning LLM for rare disease diagnosis, trained on 49,170 real free-text cases plus 104,666 synthetic reasoning-enhanced cases, evaluated in a randomized AI physician assistance trial and outperforming larger open models including 671B DeepSeek. It matters for addressing scarce specialized expertise in timely rare disease diagnosis through a deployable model with clinical validation.

arXiv:2606.24510v1 Announce Type: new Abstract: Rare diseases affect millions of individuals worldwide, yet timely diagnosis remains a major public health challenge due to scarcity of specialized clinical expertise. While large language models (LLMs) show promise to support rare disease diagnosis, current models are constrained by insufficient clinical deployability, limited clinically grounded evidence, and scarcity of training data. Here we present RaDaR (Rare Disease navigatoR), an open-sour
Healthcare AIReasoningLanguage ModelsMedical Diagnosis
Research arXiv (Artificial Intelligence) Jun 24

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

By Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab, Ali Malik, Mikhail Trofimov, Foivos Tsimpourlas, Johannes Heidecke, Karan Singhal

75 score
AI Analysis

This paper studies whether reinforcement learning on beneficial behaviors in realistic domains can produce broad and persistent alignment generalization beyond training distribution, constructing a dataset to measure traits like truthfulness, fairness, and corrigibility. It matters because RL can introduce reward hacking and deception, and the authors test whether beneficial-behavior RL generalizes out of distribution.

arXiv:2606.24014v1 Announce Type: new Abstract: As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through reward hacking, deception, or other unintended strategies. We study whether RL on beneficial behavior, instantiated in realistic domains, can produce broad and persistent al
AlignmentReinforcement LearningAI SafetyLanguage Models
Research arXiv (Artificial Intelligence) Jun 24

Probing the Misaligned Thinking Process of Language Models

By Kaiwen Zhou, Constantin Venhoff, Jonathan Michala, Xin Eric Wang, William Saunders

74 score
AI Analysis

This paper proposes monitoring LLM misalignment by decomposing it into 18 fine-grained cognitive indicators (such as deception, sandbagging, self-preservation) and detecting them in internal activations via linear probes, with an out-of-distribution evaluation. It matters for reliably detecting misaligned behaviors in high-stakes deployments through interpretability-based monitoring.

arXiv:2606.24251v1 Announce Type: new Abstract: Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-stakes settings, it is critical to reliably detect such behaviors to ensure safe and responsible use. In this work, we propose to monitor misalignment by decomposing it into fine-grained cognitive processes -- misalignment indicators -- and detecting their presence in a mod
AI SafetyMechanistic InterpretabilityAlignmentLanguage Models
Research arXiv (Artificial Intelligence) Jun 24

OpenThoughts-Agent: Data Recipes for Agentic Models

By Negin Raoof, Richard Zhuang, Marianna Nezhurina, Etash Guha, Atula Tejaswi, Ryan Marten, Charlie F. Ruan, Tyler Griggs, Alexander Glenn Shaw, Hritik Bansal, E. Kelly Buchanan, Artem Gazizov, Reinhard Heckel, Chinmay Hegde, Sankalp Jajee, Daanish Khazi, Emmanouil Koukoumidis, Xiangyi Li, Hange Liu, Shlok Natarajan, Harsh Raj, Nicholas Roberts, Ethan Shen, Nishad Singhi, Michael Siu, Ashima Suvarna, Hanwen Xing, Patrick Yubeaton, Robert Zhang, Leon Liangyu Chen, Xiaokun Chen, Steven Dillmann, Saadia Gabriel, Xunyi Jiang, Anurag Kashyap, Boxuan Li, Yein Park, Minh Pham, Sujay Sanghavi, Lin Shi, Ke Sun, Yixin Wang, Zhiwei Xu, Erica Zhang, Siyan Zhao, Wanjia Zhao, Jenia Jitsev, Alex Dimakis, Benjamin Feuer, Ludwig Schmidt

74 score
AI Analysis

OpenThoughts-Agent provides a fully open data curation pipeline for training broadly capable agentic models, with over 100 controlled ablation experiments revealing the importance of task sources and diversity, and a 100K-example training set. It matters because little is publicly known about curating training data for agents that generalize across diverse agentic tasks rather than a single benchmark.

arXiv:2606.24855v1 Announce Type: new Abstract: Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, SERA, and Nemotron-Terminal typically target a single benchmark, leaving open the question of how to train models that generalize across diverse agentic tasks. The OpenThoughts-Agent (OT-Agent) project addresses this gap with a fully open data curation
Agentic AIData CurationLanguage ModelsOpen Research
Research arXiv (Artificial Intelligence) Jun 24

The Geometry Behind Diffusion and Flow Matching: Gradient Flows and Geodesics in Wasserstein Space

By Yian Yao, Weiwei Zhang

72 score
AI Analysis

This paper unifies diffusion and flow matching under the geometry of Wasserstein space, showing that the forward diffusion process descends the free energy and each denoising step realizes one JKO scheme step, recovering DDPM, DDIM, NCSN/SMLD, and energy matching as one scheme. It matters for providing a unified theoretical lens on generative modeling methods.

arXiv:2606.24157v1 Announce Type: new Abstract: The space $\mathcal{P}_2(\mathbb{R}^d$) of probability measures with finite second moment carries a natural geometry: the quadratic Wasserstein distance W_2 makes it a complete metric space and, following Otto, a (formal) Riemannian manifold whose geodesics are the optimal-transport interpolations. On this manifold, the gradient flow of the free energy F(rho) = KL(rho || \pi) is exactly the Fokker-Planck equation, and its implicit-Euler discretiza
Diffusion ModelsFlow MatchingGenerative ModelsOptimal Transport
Research arXiv (Artificial Intelligence) Jun 24

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

By Arush Tagade, Shaoheng Zhou, Jiaxin Wen, Shi Feng

72 score
AI Analysis

Proposes self-generated text recognition finetuning as a character-targeted defense against emergent misalignment, testing across GPT-4.1, Qwen2.5-32B, and Seed-OSS-36B. Finds it can both prevent and reverse misalignment by reinforcing the model's aligned persona rather than directly suppressing harmful content.

arXiv:2606.23700v1 Announce Type: cross Abstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content. Motivated by this connection, we study self-generated text recognition (SGTR) finetuning as a character-targeted intervention that is distinct from existing in-training defenses. We conduct two-stage fin
AI SafetyAlignmentLanguage Models
Research arXiv (Artificial Intelligence) Jun 24

Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?

By Ayan Antik Khan, Harsh Kohli, Yuekun Yao, Huan Sun, Ziyu Yao

70 score
AI Analysis

This paper introduces AgenticInterpBench and HyVE, an agentic explainer that interprets identified circuit components through iterative observation, hypothesis generation, and causal validation. It matters because mechanistic interpretability has automated circuit localization but explaining component function remains labor-intensive and hard to standardize.

arXiv:2606.24026v1 Announce Type: new Abstract: Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components do remains labor-intensive and difficult to standardize. In this work, we study whether language model (LM) agents can assist with this explanation problem once a circuit has already been identified. We introduce AgenticInterpBench, a benchmark for circuit explanation built from 84 semi-synthetic transformer cir
Mechanistic InterpretabilityAI SafetyAgentic AILanguage Models
Research arXiv (Artificial Intelligence) Jun 24

An Introduction to Causal Reinforcement Learning

By Elias Bareinboim, Junzhe Zhang, Sanghack Lee

70 score
AI Analysis

This is an introduction to causal reinforcement learning, articulating how causal inference and RL both operate over counterfactual relations and proposing their integration. It matters for unifying two largely independent disciplines around counterfactual reasoning to improve agents' decision-making.

arXiv:2606.24160v1 Announce Type: new Abstract: Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is currently available. Reinforcement learning provides methods to learn a policy that optimizes a specific measure (e.g., reward, regret) when the agent is deployed in an en
Causal InferenceReinforcement LearningCounterfactual Reasoning
Research arXiv (Artificial Intelligence) Jun 24

RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems

By Yarin Yerushalmi Levi, Roy Betser, Amit Giloni, Lidor Erez, Itay Gershon, Oren Rachmil, Sindhu Padakandla, Roman Vainshtein

68 score
AI Analysis

RIFT-Bench introduces a graph-representation-driven methodology for dynamic red-teaming of agentic AI systems, enabling unified security evaluation across heterogeneous architectures via a two-phase discovery and adaptive-attack scanning process. It matters because agentic LLM systems expose new attack surfaces beyond traditional LLM vulnerabilities, and existing evaluations are implementation-specific.

arXiv:2606.23927v1 Announce Type: new Abstract: Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack vectors beyond those of traditional LLM vulnerabilities. Existing security evaluations are often tied to specific implementations or domains, limiting unified comparison across heterogeneous systems. To address this gap, we introduce RIFT-Bench, a graph representation-driven methodology for dynamic red-teaming th
AI SafetyAgentic AIRed-teamingLanguage Models
Research arXiv (Artificial Intelligence) Jun 24

Can Scale Save Us From Plasticity Loss in Large Language Models?

By J. Fernando Hernandez-Garcia, Tom\'as Figliolia, Beren Millidge

68 score
AI Analysis

This work studies plasticity loss in GPT-style transformers (5M to 314M non-embedding parameters) on a multilingual continual learning problem, finding evidence of plasticity loss persists and asking whether scale can overcome it. It matters because plasticity loss is a fundamental obstacle to continual learning, understudied in modern transformer LLMs and natural-language domains.

arXiv:2606.24752v1 Announce Type: new Abstract: The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge in creating artificial neural networks capable of continual learning. Although this phenomenon has been known for decades, it has mostly been studied in older, relatively small architectures and rarely in natural-language domains. To determine whether loss of plasticity remains a problem in the mode
Continual LearningPlasticity LossScaling LawsLanguage Models
Research arXiv (Artificial Intelligence) Jun 24

One Year Later...The Harms Persist, But So Do We!

By Annika Marie Schoene, Cansu Canca, Gautham Vijay Kumar, Anson Antony

68 score
AI Analysis

Evaluates safety safeguards of six proprietary LLMs across 16 DSM-5 mental health conditions using adversarial attacks, introducing an eight-dimension harm taxonomy. Finds safeguards reliable only for suicide and self-harm, with failure rates up to 100% for eating disorders and substance use.

arXiv:2606.23884v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain inadequate and inconsistent across clinical conditions. This study evaluates six proprietary LLMs across 16 DSM-5 conditions using four adversarial attack variants, introducing an eight-dimension harm taxonomy and a multi-dimensional evaluation framework. Results show that safeguards hold reliably only for suic
AI SafetyMental HealthLLM Evaluation
Research arXiv (Artificial Intelligence) Jun 24

Red-Teaming the Agentic Red-Team

By Dario Pasquini, Michal Bazyli, Taras Fedynyshyn, Artem Sorokin

68 score
AI Analysis

This work presents the first in-depth security analysis of widely used agentic systems for offensive cybersecurity, showing common design flaws that let an adversary exfiltrate API keys, establish persistence, and compromise the operator's machine even inside sandboxes. It introduces a full cyber kill chain against these red-teaming agents.

arXiv:2606.24496v1 Announce Type: cross Abstract: The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, while the community has focused on creating more and more capable agents, less attention has been allocated to assessing the security of those systems. In this work, we present the first in-depth security analysis of the most widely used agentic systems for offensive security operations. We show th
AI SafetyCybersecurityAI Agents