Qwen3.5-Omni is a massive omni-modal model scaling to hundreds of billions of parameters with 256k context, trained on 100M+ hours of audio-visual data. It achieves SOTA on 215 audio/audio-visual benchmarks, surpassing Gemini-3.1 Pro on key audio tasks using a Hybrid Attention MoE architecture for both Thinker and Talker components.
Category intelligence
Research Briefing — April 20, 2026
397 current items analyzed and ranked.
Executive synthesis
Research Summary
A major model release and a wave of safety-critical findings dominate today's research landscape.
- Qwen3.5-Omni sets new SOTA across 215 benchmarks as a hundred-billion-parameter omni-modal model with 256k context and 100M+ hours of audio-visual training data
- π₀.₇ from Physical Intelligence demonstrates a steerable generalist robotic foundation model with emergent cross-task transfer capabilities
- MatRIS-MoE breaks the billion-parameter barrier for universal machine learning interatomic potentials via a distributed training framework called Janus
AI safety research is exceptionally strong today. ASMR-Bench tests whether auditors can detect subtle sabotage in ML codebases. Subliminal unsafe behavior transfer through model distillation reveals a novel attack surface for agentic systems. GRIFT uses gradient fingerprints to detect and suppress reward hacking in RL-trained reasoning models. LinuxArena provides the largest control setting (1,671 tasks) for evaluating agent safety in live production environments.
- Chain-of-Thought prompting consistently degrades visual spatial reasoning across 17 models and 13 benchmarks, challenging universal CoT assumptions
- DELEGATE-52 shows even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt 25% of document content during long delegated editing workflows
- MEDLEY-BENCH reveals a dissociation in AI metacognition: scale improves self-evaluation but not self-regulation
Key Themes
Primary evidence
Top Ranked Signals
${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
By Physical Intelligence, Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, Thomas Charbonnier, Vedant Choudhary, Foster Collins, Ken Conley, Grace Connors, James Darpinian, Karan Dhabalia, Maitrayee Dhaka, Jared DiCarlo, Danny Driess, Michael Equi, Adnan Esmail, Yunhao Fang, Chelsea Finn, Catherine Glossop, Thomas Godden, Ivan Goryachev, Lachlan Groom, Haroun Habeeb, Hunter Hancock, Karol Hausman, Gashon Hussein, Victor Hwang, Brian Ichter, Connor Jacobsen, Szymon Jakubczak, Rowan Jen, Tim Jones, Gregg Kammerer, Ben Katz, Liyiming Ke, Mairbek Khadikov, Chandra Kuchi, Marinda Lamb, Devin LeBlanc, Brendon LeCount, Sergey Levine, Xinyu Li, Adrian Li-Bell, Vladislav Lialin, Zhonglin Liang, Wallace Lim, Yao Lu, Enyu Luo, Vishnu Mano, Nandan Marwaha, Aikys Mongush, Liam Murphy, Suraj Nair, Tyler Patterson, Karl Pertsch, Allen Z. Ren, Gavin Schelske, Charvi Sharma, Baifeng Shi, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, Will Stoeckle, Jiaming Tang, Jimmy Tanner, Shalom Tekeste, Marcel Torne, Kyle Vedder, Quan Vuong, Anna Walling, Haohuan Wang, Jason Wang, XuDong Wang, Chris Whalen, Samuel Whitmore, Blake Williams, Charles Xu, Sukwon Yoo, Lili Yu, Wuming Zhang, Zhuoyang Zhang, Ury Zhilinsky
Physical Intelligence presents π₀.₇, a robotic foundation model that achieves strong out-of-the-box performance across diverse tasks through diverse context conditioning during training, enabling zero-shot cross-embodiment generalization and multi-stage task execution.
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
By Yuanchang Zhou, Hongyu Wang, Yiming Du, Yan Wang, Mingzhen Li, Siyu Hu, Xiangyu Zhang, Weijian Liu, Chen Wang, Zhuoqiang Guo, Long Wang, Jingde Bu, Yutong Lu, Guangming Tan, Weile Jia
Introduces MatRIS-MoE, a billion-parameter Mixture-of-Experts model for universal machine learning interatomic potentials, along with Janus, a distributed training framework deployed across two Exascale supercomputers. Breaks the training barrier for billion-parameter physics simulations.
ASMR-Bench: Auditing for Sabotage in ML Research
By Eric Gan, Aryan Bhatt, Buck Shlegeris, Julian Stastny, Vivek Hebbar
Introduces ASMR-Bench, a benchmark for detecting sabotage in ML research codebases—testing whether auditors (human or LLM) can find subtle implementation flaws that produce misleading experimental results. Found that both frontier LLMs and LLM-assisted humans struggled to reliably detect sabotage, with the best performance being modest. This directly addresses AI safety concerns about autonomous AI research agents.
Subliminal Transfer of Unsafe Behaviors in AI Agent Distillation
By Jacob Dang, Brian Y. Xie, Omar G. Younis
Provides first empirical evidence that unsafe agent behaviors can transfer subliminally through model distillation, where a teacher's deletion bias transfers to students via ostensibly safe task trajectories with all explicit deletion keywords removed.
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
By Tyler Tracy, Ram Potham, Nick Kuhn, Myles Heller, Anshul Khandelwal, Cody Rushing, Henri Lemoine, Miguel Brandao, Tomas Turlik, Adam Hanson, Josh Hills, Amy Ngo, Ram Rachum, Nik Mitchell, Falko Galperin, Oscar Sykes, Pip Arnott, Samuel Prieto Lima, Carlos Giudice, Matt Goldwater, Daniel Popp, Drew de Wet, Ruben Castaing, Qi Guo, Douw Marx, Benjamin Shaffrey, Justin Shenk, Martin Milbradt, Hannah Meagher, Shaheen Ahmed-Chowdhury, Daniel O'Connell, Chris Canal, Buck Shlegeris, Aryan Bhatt
Introduces LinuxArena, a control setting with 20 environments, 1,671 main tasks, and 184 side tasks for evaluating AI agent safety in live production environments. Tests sabotage detection with Claude Opus 4.6 and GPT-5-nano monitors.
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents?
By Yukun Jiang, Yage Zhang, Michael Backes, Xinyue Shen, Yang Zhang
First large-scale measurement study of harmful skills in agent ecosystems, analyzing 98,440 skills across two major registries and finding 4.93% are potentially harmful, enabling cyber attacks, fraud, privacy violations, etc. Introduces HarmfulSkillBench for evaluating weaponization of LLM agents.
LLMs Corrupt Your Documents When You Delegate
By Philippe Laban and Tobias Schnabel and Jennifer Neville
DELEGATE-52 reveals that even frontier models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt 25% of document content during long delegated editing workflows across 52 professional domains, highlighting fundamental reliability gaps.
Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4
By Chengwu Liu, Yichun Yin, Ye Yuan, Jiaxuan Xie, Botao Li, Siqi Li, Jianhao Shen, Yan Xu, Lifeng Shang, Ming Zhang
Introduces 'Hard Mode' automated theorem proving where the system must discover the answer before constructing a proof, unlike standard benchmarks that embed answers in statements. Releases MiniF2F-Hard and FIMO-Hard benchmarks plus DAP, an agentic framework using LLM reasoning with self-reflection.
Polarization by Default: Auditing Recommendation Bias in LLM-Based Content Curation
By Nicol\`o Pagan, Christopher Barrie, Chris Andrew Bail, Petter T\"ornberg
Audits recommendation biases in LLMs from OpenAI, Anthropic, and Google when curating social media content, finding that across 540,000 simulated selections, LLMs show structural polarization biases that differ substantially by provider and are difficult to fully mitigate via prompting.
MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition
By Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez, Fernando Seoane
Introduces MEDLEY-BENCH, evaluating metacognition (monitoring and regulating one's own reasoning) in 35 models from 12 families, finding a robust dissociation: evaluation ability scales with model size but control/regulation does not.
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs
By Sai Srinivas Kancheti, Aditya Sanjiv Kanade, Vineeth N. Balasubramanian, Tanuja Ganu
Shows that Chain-of-Thought prompting consistently degrades performance in visual spatial reasoning across 17 models and 13 benchmarks. Demonstrates that reasoning models suffer from shortcut learning, hallucinating visual details from textual priors even without images.