Top Topic
Daily AI intelligence
Daily AI Briefing — June 3, 2026
1792 current signals analyzed across AI news, research, social media, and open-source projects.
Daily synthesis
Executive Summary
Top Story
Microsoft debuted MAI-Thinking-1, its first in-house advanced reasoning model, at Build 2026, part of a broader launch of seven new MAI models that The Verge described as medium-sized and matching leading models on software-engineering benchmarks.
Key Developments
- Alibaba/Qwen: Shipped Qwen3.7-Plus with vision, deep reasoning, and agentic tool use on the Bailian platform, while JetBrains open-sourced Mellum2, a 12B MoE coding model.
- OpenAI: Launched Codex Sites, which turns ideas into shareable apps, and expanded Codex plugins to 62 apps and 110 skills to make the agent a role specialist across professions.
- Alphabet: Shares fell after announcing an $80bn equity raise for AI infrastructure, with Berkshire Hathaway committing $10 billion to the buildout.
- NVIDIA: Released technical details and Hugging Face variants for Cosmos 3, with related work including OmniDreams for autonomous-vehicle simulation (continuing from yesterday's GTC Taipei reveal).
- r/LocalLLaMA: A two-week experiment replacing Claude with a local Qwen3.6-27B in a multi-agent orchestrator drew attention, alongside Intel Arc B70 Pro benchmarks (977 tk/s, 262k context).
Safety & Regulation
- Trump signed an executive order creating a voluntary framework for federal prerelease review of frontier models, narrowed after industry objections, which Anthropic publicly welcomed.
- 16 researchers, backed by the International Mathematical Union, published the Leiden Declaration warning of AI's encroachment on mathematics.
- NeurIPS 2026 announced its Position Paper Track will require substantially human-written submissions, prompting debate on AI authorship norms.
Research Highlights
- A neuro-AI study claimed backpropagation destroys V1 brain alignment in a single epoch, contrasting with feedback alignment, predictive coding, and STDP.
- MiniMax detailed its sparse attention (MSA) approach for native 1M-token scaling.
- Ethan Mollick highlighted a study where Gemini 2.5 beat law professors on office-hours questions with a 75% win rate, and a GitHub-data paper showing coding agents multiply output 2.2x–17.3x while human review remains the release bottleneck.
Looking Ahead
Watch whether Microsoft's move toward proprietary frontier reasoning, combined with Alphabet's $80bn raise and renewed bubble skepticism, sharpens the divide between capital-intensive scaling and the local open-model push gaining traction among practitioners.
Cross-category signals
Top Topics
Top Topic
AI Economics and Mega-Funding Skepticism
Top Topic
OpenAI Codex for White-Collar Work
Top Topic
Microsoft MAI Reasoning Models
Top Topic
US Frontier Model Governance
Top Topic
Local and Open Agentic Models
Current evidence
AI News
Frontier model and hardware releases dominated the day. NVIDIA launched Cosmos 3, a SOTA open-weights multimodal world model, alongside Nemotron 3 Ultra and RTX Spark chips. Microsoft debuted MAI-Thinking-1, its first in-house advanced reasoning model, at Build 2026. Alibaba's Qwen shipped Qwen3.7-Plus with vision, deep reasoning, and agentic tool use, while JetBrains open-sourced the Mellum2 12B MoE coding model.
- US governance: Trump signed an executive order creating a voluntary framework for federal prerelease review of frontier models, narrowed after industry objections.
- Profession risk: 16 researchers, backed by the International Mathematical Union, published the Leiden Declaration warning of AI's encroachment on mathematics.
Building on yesterday's Social unveiling of RTX Spark, now part of the broader Nvidia roundup, NVIDIA launched Cosmos 3, a Mixture-of-Transformers world model unifying language, image, video, audio, and action with autoregressive reasoner and diffusion generator towers, claiming new SOTA open-weights image and video generation. It also unveiled Nemotron 3 Ultra, a 550B-A55B open-weights LLM billed as the new US SOTA, plus the RTX Spark personal supercomputer.
Microsoft announced MAI-Thinking-1, its first advanced in-house reasoning model, described as medium-sized and matching leading models on key software-engineering benchmarks. Microsoft says it was trained from scratch on clean data without distillation from third-party models, as it loosens ties with OpenAI.
Trump signs executive order seeking early access to new AI releases
By Sanya Mansoor
Trump signed an executive order creating a voluntary framework for the federal government to vet powerful new AI models before public release, focused on cybersecurity and national security risks. The voluntary nature signals continued reluctance to impose binding rules.
Alphabet’s shares drop after announcing $80bn share sale, as AI threatens to drive up youth unemployment – as it happened
By Graeme Wearden
Alphabet's shares dropped after announcing an unprecedented $80bn equity raise to fund AI infrastructure, while Anthropic confidentially filed for a US IPO. The roundup also notes warnings that AI could drive up youth unemployment.
Warren Buffett's Berkshire Hathaway bets $10 billion on Alphabet's AI infrastructure buildout
By Matthias Bastian
Berkshire Hathaway is investing $10 billion in Alphabet's $80 billion AI infrastructure raise, with Alphabet expecting 2026 capital spending to reach $190 billion. The investment marks a notable bet on the AI buildout.
Current evidence
Research
NVIDIA dominates with two major world-model contributions: Cosmos 3 introduces an omnimodal architecture unifying multiple modalities for physical AI, while OmniDreams demonstrates real-time closed-loop AV simulation built on the Cosmos diffusion backbone.
- LEAP from Google DeepMind achieves SOTA on formal theorem proving in Lean via agentic decomposition, introducing an IMO-level benchmark
- WUSH derives closed-form transforms for joint weight-activation LLM quantization, advancing practical deployment efficiency
- Reasoning Primitive Induction distills reusable reasoning moves from agent traces, surprisingly outperforming the source agent across tasks
- TRAP exposes a novel attack surface on Vision-Language-Action robotic models by hijacking Chain-of-Thought reasoning via adversarial patches
Safety and alignment research features prominently: a principled cybersecurity refusal framework (Kolter et al.) addresses agent deployment boundaries, while a position paper argues solipsistic superintelligence is unlikely to be cooperative. Interpretability advances include a graph-based reasoning structure benchmark and a causal geometric decomposition of how prompting steers internal LLM representations.
Cosmos 3: Omnimodal World Models for Physical AI
By Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson, Carlos Casanova, Ting-Yun Chang, Yan Chang, Yu-Wei Chao, Prithvijit Chattopadhyay, Roshan Chaudhari, Chieh-Yun Chen, Junyu Chen, Ke Chen, Qizhi Chen, Wenkai Chen, Xiaotong Chen, Yu Chen, An-Chieh Cheng, Click Cheng, Xiu Chia, Jeana Choi, Chaeyeon Chung, Wenyan Cong, Yin Cui, Magdalena Dadela, Nalin Dadhich, Wenliang Dai, Joyjit Daw, Alperen Degirmenci, Rodrigo Vieira Del Monte, Robert Denomme, Sameer Dharur, Marco Di Lucca, Ke Ding, Wenhao Ding, Yifan Ding, Yuzhu Dong, Nicole Drumheller, Yilun Du, Aigul Dzhumamuratova, Aleksandr Efitorov, Hamid Eghbalzadeh, Naomi Eigbe, Imad El Hanafi, Hassan Eslami, Benedikt Falk, Jiaojiao Fan, Jim Fan, Amol Fasale, Sergiy Fefilatyev, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Vikram Fugro, Prashant Gaikwad, TJ Galda, Katelyn Gao, Yihuai Gao, Wenhang Ge, Sreyan Ghosh, Arushi Goel, Vivek Goel, Akash Gokul, Rama Govindaraju, Jinwei Gu, Miguel Guerrero, Elfie Guo, Aryaman Gupta, Siddharth Gururani, Hugo Hadfield, Song Han, Ankur Handa, Zekun Hao, Mohammad Harrim, Ali Hassani, Nathan Hayes-Roth, Yufan He, Chris Helvig, Cyrus Hogg, Madison Huang, Michael Huang, Sophia Huang, Yufan Huang, Jacob Huffman, DeLesley Hutchins, Suneel Indupuru, Boris Ivanovic, Arihant Jain, Joel Jang, Ryan Ji, Yanan Jian, Dongfu Jiang, Jingyi Jin, Atharva Joshi, Nikhilesh Joshi, Pranjali Joshi, Jaehun Jung, Weiwei Kang, Scott Kassekert, Jan Kautz, Ashna Khetan, Julia Kiczka, Slawek Kierat, Gwanghyun Kim, Kuno Kim, Sunny Kim, Kezhi Kong, Xin Kong, Zhifeng Kong, Tomasz Kornuta, Egor Krivov, Hui Kuang, Saurav Kumar, Chia-Wen Kuo, George Kurian, Wojciech Kutak, JF Lafleche, Himangshu Lahkar, Omar Laymoun, Jayjun Lee, Sanggil Lee, Gabriele Leone, Boyi Li, Freya Li, Jiajun Li, Jinfeng Li, Ling Li, Pengcheng Li, Shangru Li, Tingle Li, Xiaolong Li, Xuan Li, Zhaoshuo Li, Zhiqi Li, Hao Liang, Maosheng Liao, Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Sifei Liu, Zihan Liu, Hai Loc Lu, Xiangyu Lu, Alice Luo, Ruipu Luo, Wenjie Luo, Jiangran Lyu, Martin Ding Ma, Nic Ma, Qianli Ma, Dawid Majchrowski, Louis Marcoux, Miguel Martin, Qing Miao, Ashkan Mirzaei, Shreyas Misra, Kaichun Mo, Durra Mohsin, Hyejin Moon, Pawel Morkisz, Saeid Motiian, Kirill Motkov, Seungjun Nah, Yashraj Narang, Deepak Narayanan, Thabang Ngazimbi, Julian Ouyang, David Page, Yatian Pang, Sehwi Park, Mahesh Patekar, Mostofa Patwary, Marco Pavone, Trung Pham, Wei Ping, Soha Pouya, Shrimai Prabhumoye, Varun Praveen, Delin Qu, Hesam Rabeti, Morteza Ramezanali, Marilyn Reeb, Xuanchi Ren, Kristen Rumley, Wojciech Rymer, Jun Saito, Yeongho Seol, John Shao, Piyush Shekdar, Tianwei Shen, Humphrey Shi, Min Shi, Stella Shi, Kevin Shih, Mohammad Shoeybi, Mateusz Sieniawski, Shuran Song, Alexander Sotelo, Amir Sotoodeh, Sunil Srinivasa, Vignesh Srinivasakumar, Bartosz Stefaniak, Rahul Heinrich Steiger, Shangkun Sun, Jiaxiang Tang, Shitao Tang, Yangyang Tang, Yue Tang, Tolou Tavakkoli, Kayley Ting, Krzysztof Tomala, Wei-Cheng Tseng, Jibin Varghese, Sergei Vasilev, Thomas Volk, Raju Wagwani, Roger Waleffe, Andrew Z. Wang, Boxiang Wang, Haoxiang Wang, Qiao Wang, Shihao Wang, Shijie Wang, Ting-Chun Wang, Yan Wang, Yu Wang, David Wehr, Fangyin Wei, Xinshuo Weng, Jay Zhangjie Wu, Kedi Wu, Hongchi Xia, Summer Xiao, Tianjun Xiao, Kevin Xie, Daguang Xu, Jiashu Xu, Mengyao Xu, Ruqing Xu, Xingqian Xu, Yao Xu, Dinghao Yang, Dong Yang, Hans Yang, Xiaodong Yang, Xuning Yang, Yichu Yang, Yurong You, Zhiding Yu, Hao Yuan, Simon Yuen, Xiaohui Zeng, Pengcuo Zeren, Cindy Zha, Haotian Zhang, Jenny Zhang, Jing Zhang, Liangkai Zhang, Paris Zhang, Shun Zhang, Xuanmeng Zhang, Zhizheng Zhang, Ann Zhao, Yilin Zhao, Yuliya Zhautouskaya, Charles Zhou, Fengzhe Zhou, Shilin Zhu, Yuke Zhu, Dima Zhylko, Artur Zolkowski
Building on yesterday's News announcement, here's the full Cosmos 3 technical paper, Cosmos 3 from NVIDIA is a family of omnimodal world models that jointly process and generate language, image, video, audio, and action within a unified mixture-of-transformers architecture, subsuming vision-language models, video generators, world simulators, and world-action models. It claims state-of-the-art across understanding and generation tasks for Physical AI.
NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation
By NVIDIA, :, Aarti Basant, Amlan Kar, Despoina Paschalidou, Fangyin Wei, Francesco Ferroni, Guillermo Garcia Cobo, Haithem Turki, Huan Ling, Jaewoo Seo, James Lucas, Jay Zhangjie Wu, Jialiang Wang, Jonathan Lorraine, Jun Gao, Kai He, Katarina Tothova, Kevin Xie, Micha{\l} Tyszkiewicz, Qi Wu, Riccardo de Lutio, Ruilong Li, Sanja Fidler, Seung Wook Kim, Tianchang Shen, Tianshi Cao, Tobias Pfaff, William Lew, Xindi Wu, Xuanchi Ren, Yifan Lu, Yuxuan Zhang, Zan Gojcic, Zian Wang
NVIDIA's OmniDreams is a foundation generative world model, post-trained from the Cosmos diffusion model, for real-time closed-loop autonomous vehicle simulation that autoregressively generates action-conditioned sensor observations. Addresses long-tail scenario evaluation beyond reconstruction-based simulators.
LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks
By Po-Nien Kung, Linfeng Song, Dawsen Hwang, Jinsung Yoon, Chun-Liang Li, Simone Severini, Mirek Ol\v{s}\'ak, Edward Lockhart, Quoc V Le, Burak Gokturk, Thang Luong, Tomas Pfister, Nanyun Peng
LEAP is an agentic framework enabling general-purpose foundation models to achieve state-of-the-art automated formal theorem proving in Lean by decomposing problems and iterating with the Lean compiler. It introduces Lean-IMO-Bench for rigorous evaluation.
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
By Jiale Chen, Vage Egiazarian, Roberto L. Castro, Torsten Hoefler, Dan Alistarh
Derives closed-form near-optimal linear blockwise transforms for joint weight-activation LLM quantization, called WUSH, combining a Hadamard backbone with a data-dependent second-moment component. It provides provable near-optimality for both integer and floating-point quantizers.
A New Framework for Cybersecurity Refusals in AI Agents
By Eliot Krzysztof Jones, Mateusz Dziemian, Matt Fredrikson, J Zico Kolter
This paper presents the first framework for establishing refusal boundaries for AI agents in offensive cybersecurity contexts, defining principled refusal criteria, task categories warranting refusal, and an evaluation methodology under benign and adversarial conditions. It complements proficiency-focused cyber benchmarks with a safety-refusal dimension.
Current evidence
Social Media
AI economics skepticism dominated discussions today. Gary Marcus went viral arguing AI will eventually 'fall apart' financially, citing commodity tech, no moats, and unsustainable capex. Timnit Gebru echoed criticism of consulting-firm hype cycles.
- OpenAI drove major product news with Codex Sites (turning ideas into shareable apps) and expanded Codex plugins turning the agent into a role specialist across 62 apps and 110 skills.
- Ethan Mollick anchored the research conversation, highlighting a study where Gemini 2.5 beat law professors on office-hours questions (75% win rate), and a GitHub-data paper showing coding agents multiply output (2.2x–17.3x) while human review remains the release bottleneck.
- Google DeepMind promoted its Co-Scientist multi-agent hypothesis system, expanding access via Gemini for Science.
Policy and institutional moves rounded out the day. Anthropic welcomed a US AI Executive Order and expanded Project Glasswing / Claude Mythos Preview to ~150 more organizations. NeurIPS 2026 announced its Position Paper Track will require substantially human-written submissions, sparking debate on AI authorship norms, while a notable open-model researcher announced their departure from Ai2.
Why things will eventually fall apart: 1. Everybody, even Google, seems to be treating AI as if it ...
By @GaryMarcus
Gary Marcus lays out a five-point thesis that AI will eventually fall apart financially because everyone builds the same commodity tech with no moat, preventing monopoly pricing and forcing price wars that make returns modest relative to spending.
Building apps has never been easier. With Sites, Codex can turn your work, ideas, and plans into an...
By @OpenAI
OpenAI launches Codex Sites, letting Codex turn work, ideas, and plans into shareable interactive websites or apps, rolling out to Business and Enterprise plans.
Law professors wrote questions they were asked during office hours. Gemini 2.5 & humans answered...
By @emollick
Ethan Mollick describes a study where Gemini 2.5 answered law professors' office-hours questions and beat human professors with a 75 percent win rate while being rated less harmful, with newer models doing even better.
This Executive Order is an important step in strengthening America’s leadership in AI. We look for...
By @AnthropicAI
Anthropic publicly welcomes a US Executive Order on AI and signals intent to collaborate with the White House on implementation.
Big paper on AI coding agents using Github & other data The auto-complete tools (Copilot) led ...
By @emollick
Ethan Mollick summarizes a big paper using GitHub data showing autocomplete tools led to 2.2x more code, local agents 7.4x, and remote coding agents 17.3x, but human bottlenecks meant releases only rose 30 percent.