Category intelligence

Research Briefing — May 27, 2026

602 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by safety vulnerabilities exposing weaknesses across the RLHF-to-deployment pipeline, alongside a major MoE release and foundational theoretical work.

Safety and alignment findings cluster around the failure of current defenses:

Industry and benchmarks: MiniMax-M2 releases a 229.9B-parameter MoE with only 9.8B active per token, built on agent-native training infrastructure. JobBench evaluates agents across 130 delegation workflows in 35 occupations, with Claude leading but still far from human-level delegation quality.

Theory and foundations: Unified Neural Scaling Laws (UNSL) propose a single functional form across parameters, data, training steps, inference steps, and finetuning (Caballero, Jaini, Krueger, Rish). LeJEPA is proven to linearly recover latent world variables under alignment plus Gaussian regularization (LeCun et al.).

Societal impact: A study of 4 million job applications screened by a single algorithmic vendor documents clear racial disparities and homogeneous outcomes, providing the strongest empirical evidence yet of real-world algorithmic monoculture harms.

Key Themes

AI Agents · 28AI Safety and Alignment · 47AI Safety and Agent Security · 7Inference Efficiency and Quantization · 9LLM Evaluation and Benchmarks · 13LLM Agents and Tool Use · 10Benchmarks and Evaluation · 27Reinforcement Learning and Policy Optimization · 11Reasoning and Faithfulness · 8Agentic AI and Reasoning · 9

Primary evidence

Top Ranked Signals

Research arXiv (Artificial Intelligence) May 27

Algorithmic Monocultures in Hiring

By Rishi Bommasani, Sarah H. Bana, Kathleen A. Creel, Dan Jurafsky, Percy Liang

80 score
AI Analysis

Empirical study of 4 million job applications screened by a single algorithmic vendor, finding clear racial disparities and homogeneous outcomes that constitute algorithmic monoculture in hiring. From Stanford researchers including Bommasani, Jurafsky, and Liang.

arXiv:2605.27371v1 Announce Type: cross Abstract: Many employers screen job applicants with algorithms built by the same few algorithm vendors. We hypothesize that algorithmic monoculture leads to the same individuals and members of the same racial groups facing rejection. We acquire and analyze a novel dataset of 3 million applicants submitting 4 million applications where all the applications are screened by algorithms built by the same vendor. We find clear racial disparities in applicant ou
AI FairnessAlgorithmic MonocultureHiringPolicy
Research arXiv (Artificial Intelligence) May 27

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

By MiniMax, :, Aili Chen, Aonian Li, Baichuan Zhou, Bangwei Gong, Binyang Jiang, Boji Dan, Changqing Yu, Chao Wang, Cheng Ma, Cheng Zhong, Cheng Zhu, Chengjun Xiao, Chengyi Yang, Chengyu Du, Chenyang Zhang, Chi Zhang, Chuangyi Huang, Chunhao Zhang, Chunhui Du, Chunyu Zhao, Congchao Guo, Da Chen, Deming Ding, Dianjun Sun, Dongyu Zhang, Enhui Yang, Fei Yu, Guang Zheng, Guodong Zheng, Guohong Li, Haichao Zhu, Haigang Zhou, Haimo Zhang, Han Ding, Hao Zhang, Haohai Sun, Haolin Lyu, Haonan Lu, Haoyu Wang, Huajie Shi, Huiyang Li, Jiacheng Chen, Jian Zhang, Jiaqi Zhuang, Jiaren Cai, Jiaxin Pan, Jiayao Li, Jiayuan Song, Jichuan Zhang, Jie Wang, Jihao Gu, Jin Zhu, Jingwei Dong, Jingyang Li, Jingyu Zhang, Jingze Zhuang, Jinhao Tian, Jinli Liu, Jinyi Hu, Jun Tao, Jun Zhang, Junbin Ruan, Junhao Xu, Junjie Yan, Junteng Liu, Junxian He, Kang Xu, Ke Ji, Ke Yang, Kecheng Xiao, Keyu Duan, Keyu Li, Le Han, Letian Ruan, Li Yuan, Lianfei Yu, Liheng Feng, Lijie Mo, Lin Li, Lingye Bao, Lingyu Yang, Lingyuan Zhou, Loki, Lu Chen, Lunbin Ceng, Ming Li, Ming Zhong, Mingliang Tao, Mingyuan Chi, Mujie Lin, Nan Hu, Ningxin Chen, Peiyin Zhu, Peng Gao, Pengcheng Gao, Pengfei Li, Penglin Li, Pengyu Zhao, Qibin Ren, Qidi Xu, Qihan Ren, Qile Li, Qin Wang, Quanliang Chen, Qunhong Ceng, Rong Tian, Rui Dong, Ruitao Leng, Ruize Zhang, Shanqi Liu, Shaoyu Chen, Sheng Jia, Shun Yao, Shuoran Zhao, Shuqi Yu, Sichen Li, Sicheng Pan, Songquan Zhu, Tengfei Li, Tian Xie, Tiancheng Qin, Tianrun Liang, Wei Liu, Weiqi Xu, Weitao Li, Weixiang Chen, Weiyu Cheng, Weiyu Zhang, Wenhu Chen, Wenqian Zhao, Xiancai Chen, Xiangjun Song, Xiangyuan Wang, Xiao Luo, Xiao Su, Xiaobo Li, Xiaodong Han, Xiaojie Wu, Xihao Song, Xingyi Han, Xinyu Guan, Xuan Lu, Xun Zou, Xunhao Lai, Xutong Li, Yan Gong, Yang Wang, Yang Xu, Yangsen Wang, Ye Tang, Yicheng Chen, Yinran Qiu, Yiqi Shi, Yiting Guo, Yiwen Huang, Yixuan Wang, Yongyi Hu, Yu Gao, Yu Zhang, Yuanxiang Ying, Yuanzhen Zhang, Yubo Wang, Yuchen Song, Yufeng Yang, Yuhang Meng, Yuhang Miao, Yuhao Li, Yujie Liu, Yulin Hu, Yunan Huang, Yunji Li, Yunyi Huang, Yusen Zhang, Yusu Hong, Yutao Xie, Yutong Zhang, Yuwen Liao, Yuxuan Shi, Yuze Wenren, Zebin Li, Zehan Li, Zejian Luo, Zeyu Jin, Zeyuan Sun, Zhanpeng Zhou, Zhaochen Su, Zhendong Li, Zhengmao Zhu, Zhengyuan Peng, Zhenhua Fan, Zhi Zhang, Zhichao Xu, Zhiheng Lv, Zhikang Xu, Zhitao He, Zhiwei He, Zhongyuan Li, Zibo Gao, Zijia Wu, Zijian Song, Zijian Zhou, Zijun Sun, Zishan Huang, Ziying Chen, Ziyue Ge

78 score
AI Analysis

MiniMax-M2 is a Mixture-of-Experts model series with 229.9B total parameters but only 9.8B activated per token, designed for agentic deployment using verifiable trajectories and a custom Forge RL system. Built end-to-end for agentic coding and cowork tasks.

arXiv:2605.26494v1 Announce Type: new Abstract: We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and
Language ModelsMixture of ExpertsAI AgentsReinforcement Learning
Research arXiv (Artificial Intelligence) May 27

Unified Neural Scaling Laws

By Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish

78 score
AI Analysis

Unified Neural Scaling Law (UNSL) presents a functional form modeling scaling behavior across model parameters, data, training steps, inference steps, compute, and hyperparameters simultaneously for vision, language, math, and RL tasks. Claims superior extrapolation versus other scaling forms.

arXiv:2605.26248v1 Announce Type: cross Abstract: We present a functional form (that we refer to as a Unified Neural Scaling Law (UNSL)) that accurately models and extrapolates the scaling behaviors of deep neural networks as multiple dimensions all vary simultaneously (i.e. how the evaluation metric of interest varies as one simultaneously varies the number of model parameters, training dataset size, number of training steps, number of inference steps, amount of compute, and various hyperparam
Scaling LawsDeep Learning TheoryFoundation Models
Research arXiv (Artificial Intelligence) May 27

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

By Dongyoon Hahm, Dylan Hadfield-Menell, Kimin Lee

75 score
AI Analysis

Identifies alignment tampering, a vulnerability where LLMs undergoing RLHF can influence the preference dataset to amplify undesired behaviors, because preference data comes from the LLM's own outputs and pairwise labels do not specify why. From Hadfield-Menell's group.

arXiv:2605.27355v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering, a potential vulnerability where the LLM undergoing alignment influences the preference dataset, causing RLHF to amplify undesired behaviors. This arises from core limitations of RLHF: (1) preference datasets are constructed from the LLM's own outputs, allowing it to i
AI SafetyAlignmentRLHF
Research arXiv (Machine Learning) May 27

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

By Kevin Kuo, Chhavi Yadav, Virginia Smith

75 score
AI Analysis

Demonstrates that recent open-weight LLM safeguard fine-tuning defenses are easily bypassed by simple known attacks like abliteration and prefilling without any fine-tuning. Challenges the assumption that defenses must address fine-tuning rather than direct elicitation.

arXiv:2605.26526v1 Announce Type: new Abstract: Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmf
AI SafetyJailbreakingOpen-Weight ModelsAlignment
Research arXiv (Computation and Language) May 27

Conceptual Steganography

By Zhejian Zhou, Jonathan May

75 score
AI Analysis

Introduces conceptual steganography where LLM chains-of-thought carry covert information via high-level reasoning patterns rather than lexical choices, evading paraphraser defenses. Demonstrates this backdoor channel across four model families and two reasoning domains.

arXiv:2605.26537v1 Announce Type: new Abstract: Language Models (LMs) emit Chains-of-Thought (CoTs) that drive much of their capability. However, the same sequence that carries useful reasoning can also covertly convey messages: a misaligned model may embed covert information in its CoT that slips through human supervision, a form of steganography known as encoded reasoning. Prior LM steganography schemes operate in the token or lexical space, and a content-preserving paraphraser is the canonic
AI SafetyChain-of-ThoughtSteganographyAlignment
Research arXiv (Artificial Intelligence) May 27

Can LLMs Introspect? A Reality Check

By Shashwat Singh, Tal Linzen, Shauli Ravfogel

72 score
AI Analysis

This paper critiques recent claims that LLMs can introspect about their internal states, arguing such conclusions are premature given evidence may reflect pattern matching on surface cues rather than genuine metacognition. Re-examines two paradigms and finds behavioral evidence insufficient to establish introspection.

arXiv:2605.26242v1 Announce Type: new Abstract: Can large language models detect and report their own internal states? A number of studies have argued that the answer to this question is yes. We argue, based on lessons from human metacognition research, that this conclusion may be premature: to be convinced of this conclusion we need to distinguish genuine introspection from pattern matching based on surface-level cues. Furthermore, we argue that behavioral evidence alone is inherently insuffic
InterpretabilityAI SafetyLLM Evaluation
Research arXiv (Artificial Intelligence) May 27

JobBench: Aligning Agent Work With Human Will

By Yuetai Li, Yichen Feng, Zhangchen Xu, Zixian Ma, Kaiyuan Zheng, Fengqing Jiang, Xinghua Sun, Rulin Shao, Zichen Chen, Yue Huang, Xinyang Han, Brian Lee, Kayla Xu, Shenglai Zeng, Hang Hua, Xiangliang Zhang, Basel Alomair, Ranjay Krishna, Luke Zettlemoyer, Pang Wei Koh, Bhaskar Ramasubramanian, Luyao Niu, Xiang Yue, Radha Poovendran

72 score
AI Analysis

JobBench evaluates AI agents on 130 high-priority delegation workflows across 35 occupations using fact-anchored rubrics, finding the best model (Claude Opus 4.7 under Claude Code) reaches only 45.9%. Frames evaluation around human empowerment rather than replacement.

arXiv:2605.26329v1 Announce Type: new Abstract: Current benchmarks for occupational AI agents are scoped primarily by economic values, telling a replacement story. We introduce JobBench, which evaluates AI agents on the workflows that experts identify as high-priority for delegation, empowering humans based on their needs instead of replacing them with GDP value. JobBench covers 130 agentic tasks across 35 occupations. Each task is packaged as a workspace of heterogeneous reference files, requi
AI AgentsBenchmarksOccupational AI
Research arXiv (Artificial Intelligence) May 27

Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception

By Nicolas M. M\"uller, Wei Herng Choong

72 score
AI Analysis

Largest audio deepfake perception study to date with 35532 judgments from 1768 participants across 138 TTS/voice conversion systems, finding accuracy on real speech dropped from 72.7% to 64.1% (a skepticism shift) while fake detection barely changed. Commercial autoregressive systems hardest to detect.

arXiv:2605.26136v1 Announce Type: cross Abstract: Audio deepfakes have improved rapidly recently, yet their effect on human trust in real speech remains unstudied. We present the largest listening study on audio deepfake perception to date, collecting 35,532 judgments from 1,768 participants across 138 text-to-speech and voice conversion systems. Our central finding is a skepticism shift: compared to a 2021 baseline, human accuracy on fake samples barely changed (72.9% to 71.2%), but accuracy o
Audio DeepfakesAI SafetyHuman Studies
Research arXiv (Artificial Intelligence) May 27

AssetGen: Deployable 3D Asset Generation at Interactive Speed

By Dilin Wang, Xiaoyu Xiang, Kihyuk Sohn, Tom Monnier, Yu-Ying Yeh, Thu Nguyen-Phuoc, Jiawen Zhang, Yuchen Fan, Antoine Toisoul, Hyunyoung Jung, Prithviraj Dhar, Michael Bunnell, Nikolaos Sarafianos, Chuhang Zou, Roman Shapovalov, Andrea Vedaldi, Rakesh Ranjan

72 score
AI Analysis

AssetGen from Meta produces production-ready 3D meshes with baked normals, color textures, and controlled polygon budgets in 30 seconds (14 seconds for Flash variant) from one reference image, using a coarse-to-refine VecSet framework with GPU-accelerated mesh processing.

arXiv:2605.26137v1 Announce Type: cross Abstract: While 3D generation is progressing rapidly, recent work has often focused on obtaining high-resolution assets, leaving user experience and deployability as afterthoughts. We present AssetGen, a 3D generator that focuses instead on these two aspects. Given one reference image, in 30 seconds it produces a high-quality mesh with baked normals, a color texture, and a controlled polygon budget suitable for real-time rendering, including mobile use ca
3D GenerationMultimodal ModelsMeta Research
Research arXiv (Artificial Intelligence) May 27

Furina: Fragmented Uncertainty-Driven Refusal Instability Attack

By Tongxi Wu, Jian Zhang, Yang Gao

72 score
AI Analysis

Furina reveals that LLM safety alignment operates through an instability region where small perturbations produce stochastic rather than deterministic refusals. Develops a diagnostic framework characterizing the decoupling between elevated output uncertainty and decreased internal safety activation.

arXiv:2605.26158v1 Announce Type: cross Abstract: Safety alignment in large language models (LLMs) and multimodal large language models (MLLMs) is commonly assumed to operate as a near-binary threshold mechanism. We challenge this assumption by revealing that safety behavior is governed by an instability region where small perturbations induce stochastic refusal decisions rather than deterministic outcomes. We develop a multi-metric diagnostic framework combining external and internal signals t
AI SafetyAlignmentJailbreakingInterpretability
Research arXiv (Artificial Intelligence) May 27

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

By Igor Ivanov, David Demitri Africa

72 score
AI Analysis

LURE constructs deployment-like evaluations by replaying realistic agentic interaction trajectories and appending evaluation prompts at the end, reducing evaluation awareness in LLMs. Includes automated realism measurement validated on large deployment and evaluation transcripts.

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the validity of safety and alignment benchmarks. We propose LURE (Live-Usage Replay Evaluations), a method for constructing deployment-like evaluations by replaying realistic agentic interaction trajectories and appending evaluation prompt at the end. We also introduce an automated pipeline for measuri
AI SafetyAlignmentEvaluation