Daily AI intelligence

Daily AI Briefing — May 25, 2026

1230 current signals analyzed across AI news, research, social media, and open-source projects.

Daily synthesis

Executive Summary

Top Story

DeepMind's AI agent autonomously solved 9 of 353 open Erdős problems at a few hundred dollars each via an LLM-Lean formal verification loop, though a competing Princeton/Swarat neurosymbolic approach reportedly matched the result in three days — fueling debate over whether scaled LLMs or neurosymbolic methods deserve credit.

Key Developments

Safety & Regulation

Research Highlights

Local Inference

Looking Ahead

With DeepSeek V4 Pro's pricing collapsing the cost floor and local inference stacks becoming daily-driver viable, the commercial pressure on closed-frontier labs will intensify just as red-teaming results show their safety guardrails can be broken at near-100% rates with modest budgets.

Cross-category signals

Top Topics

Top Topic

DeepSeek Pricing & GenAI Bubble Debate

DeepSeek V4 Pro's pricing (roughly 11-34x cheaper than GPT-5.5 and Claude Opus 4.7) reignited bubble debates across communities, with a viral r/OpenAI post arguing it 'popped the American AI bubble.' Gary Marcus likened LLM companies to airlines facing thin margins and commoditization, while Andrej Karpathy praised a large company 'calling BS' on AI hype. The Guardian piece on UK 'AI washing' added another angle on hype-driven rebranding.
3 Social 2 News

Top Topic

Coding Agents and Open-Source Friction

Microsoft Research open-sourced Webwright, a terminal-native web agent that writes Playwright code and scores 60.1% on Odysseys versus 33.5% for base GPT-5.4. Anthropic's Boris Cherny drove engagement recommending Claude Code's auto mode for parallel 'multi-clauding,' while Greg Brockman highlighted that Codex is open source. Meanwhile, vLLM publicly banned a contributor for AI-slop PRs, and a Reddit user vibe-coded a civic site that forced a Greek ministry to remove a fake tax-fraud hotline page.
4 Social 2 News

Top Topic

DeepMind's Erdős Math Breakthrough

Google DeepMind's AI agent autonomously solved 9 of 353 open Erdős problems at a few hundred dollars per problem via an LLM-Lean formal verification loop, dominating r/singularity discussion. Gary Marcus amplified a counter-narrative claiming a young Princeton professor beat OpenAI and Swarat's neurosymbolic approach at the same Erdős benchmark in three days. The dual storyline sparked debate over whether the win belongs to scaled LLMs or neurosymbolic methods.
2 Social

Top Topic

AI Safety, Red-Teaming and Trust

Multiple arXiv papers showed Test-Time Training breaks safety guardrails at 95% ASR with cross-model transfer, PoisonForge compromised 11 of 12 instruction-tuned LLMs at 1% poison budget, and cross-lingual multimodal red-teaming exposed non-uniform vulnerabilities in Claude Sonnet 4.5, GPT-5, Pixtral Large, and Qwen Omni. A separate paper traced geopolitical bias to post-training rather than pretraining. On Reddit, auditory prompt injection attacks via inaudible sounds in media emerged as a new class of voice-assistant exploit, alongside ChatGPT bias-toward-institutions concerns.

Top Topic

Local Inference and Hardware Diversification

Clément Delangue flagged that llama.cpp's MTP support makes Qwen3.6-27B 78% faster on A10G, pushing local models toward daily-driver viability. r/LocalLLaMA debated whether NVIDIA remains the default for 2026, while hipEngine brought fast native Qwen 3.6 inference to RDNA3/Strix Halo and BitCPM-CANN demonstrated native 1.58-bit training on Huawei Ascend NPUs. A 768GB Optane DIMM rig reportedly ran trillion-parameter Kimi K2.5 at ~4 tok/s.
1 Social

Top Topic

AI Backlash, Datacenters and Society

A Scottish charity analysis warned that Scotland's 2022 green datacentres' policy ignores AI emissions entirely, predating ChatGPT. A Futurology thread highlighted booed commencement speakers, blocked datacenters, and plummeting AI poll numbers in the US. Ethan Mollick predicted mass realization of AI-generated content saturating social media, blogs, and scientific papers, while Wendy Liu's Guardian opinion argued intellectual struggle is essential to being human.
3 News 1 Social

Current evidence

AI News

View category →

Research and model releases dominate the frontier signal this cycle:

  • Microsoft Research open-sourced Webwright, a terminal-native web agent that writes Playwright code instead of clicking actions, scoring 60.1% on Odysseys versus 33.5% for base GPT-5.4.
  • NVIDIA released Gated DeltaNet-2, a linear attention layer decoupling erase/write in the delta rule, beating Mamba-2, Mamba-3, and KDA at 1.3B parameters on 100B tokens.
  • StepFun launched StepAudio 2.5 Realtime, an end-to-end voice LLM with roleplay-specific RLHF, paralinguistic comprehension, and million-scale persona data augmentation.

Policy, industry, and cultural threads highlight tensions around AI's footprint and hype:

85 score
AI Analysis

Microsoft Research released Webwright, an open-source terminal-native web agent framework that lets agents write Playwright code instead of taking single browser actions. It scores 60.1% on Odysseys benchmark versus base GPT-5.4's 33.5%.

Most web agents today drive a browser one action at a time. The model receives the current page state — as a screenshot or DOM text — and predicts the next click, keypress, or scroll. This action-at-a-time design made sense when language models had limited reasoning ability. As models have become more capable at writing and debugging code, that rigid loop has become a constraint rather than a structure that helps. Microsoft Research’s AI Frontiers lab built a different approach. Their n
AI agentsOpen sourceMicrosoft Research
82 score
AI Analysis

NVIDIA released Gated DeltaNet-2, a linear attention layer that decouples erase and write operations in the delta rule via channel-wise gates. Trained at 1.3B parameters on 100B FineWeb-Edu tokens, it outperforms Mamba-2, Mamba-3, Gated DeltaNet, and KDA.

Linear attention replaces the unbounded KV cache of softmax attention with a fixed-size recurrent state. This cuts sequence mixing to linear time and decoding to constant memory. The hard part is not what to forget. It is how to edit a compressed memory without scrambling existing associations. NVIDIA has released Gated DeltaNet-2, a linear attention layer that targets that bottleneck. The model decouples the active memory edit into two channel-wise gates. It is trained at 1.3B parameters on
Model architectureLinear attentionNVIDIA research
75 score
AI Analysis

StepFun released StepAudio 2.5 Realtime, an end-to-end real-time speech LLM with customizable persona capabilities, roleplay-specific RLHF, and paralinguistic comprehension. It supports Chinese and English via a WebSocket API and uses million-scale persona data augmentation.

StepFun, the Shanghai-based AI lab, released StepAudio 2.5 Realtime. It is an end-to-end real-time speech large language model with fully customizable persona capabilities. StepAudio 2.5 Realtime is a voice model that operates in real time. Unlike pipeline-based systems that separate speech recognition, reasoning, and synthesis into sequential steps, this is an end-to-end model. Audio goes in and audio comes out through a single unified system. The model supports Chinese and English. It c
Voice AIModel releaseMultimodal
News AI (artificial intelligence) | The Guardian May 24

Scotland’s ‘green datacentres’ policy ignores emissions impact of AI, analysis shows

By Aisha Down

55 score
AI Analysis

A Scottish charity analysis warns that Scotland's 'green datacentres' policy, defined in 2022 before ChatGPT, ignores the massive carbon emissions impact of AI workloads. The policy underpins UK efforts to attract AI investment but may mask significant environmental costs.

Definition of green facilities made in 2022, before release of ChatGPT, says Action to Protect Rural ScotlandA Scottish government policy designed to encourage datacentres to build in Scotland could lead to a massive volume of carbon emissions being ignored, according to an analysis by a Scottish charity.“Green datacentres” are at the heart of Scotland’s ambitions to develop economically. Enshrined in national policy, they are part of a larger, UK-wide effort to attract big AI investment to Scot
AI infrastructureSustainabilityPolicy
News AI (artificial intelligence) | The Guardian May 24

‘We’re expanding the cinematic toolbox’: AI fault lines on show at Cannes

By Nadia Khomami Arts and culture correspondent

50 score
AI Analysis

At Cannes, divisions over generative AI in filmmaking deepened, with Darren Aronofsky championing AI through his studio Primordial Soup while Guillermo del Toro voiced strong opposition. The festival highlighted AI as Hollywood's most divisive issue.

Darren Aronofsky among proponents of using technology, while Guillermo del Toro says he would ‘rather die’Under a white marquee on Cannes’ Croisette beach, with the Mediterranean glistening behind him and superyachts drifting across the horizon, the director Darren Aronofsky addressed an audience of executives and tech evangelists gathered for an “AI for Talent” summit.“There’s so much pushback against AI,” said Aronofsky, who has faced criticism over his embrace of generative AI projects though
Generative AICreative industriesFilm

Current evidence

Research

View category →

Today's research is dominated by safety/red-teaming findings on frontier models and foundational advances in training, retrieval, and RL.

Safety, Alignment & Red-Teaming:

Foundations & Systems:

RL & Evaluation Methodology:

Research arXiv (Artificial Intelligence) May 25

Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified Systems

By Shubham Agarwal, Alexander Krentsel, Shu Liu, Mert Cemri, Audrey Cheng, Rui Meng, Tomas Pfister, Chun-Liang Li, Sylvia Ratnasamy, Aditya Parameswaran, Matei Zaharia, Ion Stoica, Mohsen Lesani

78 score
AI Analysis

Inductive Deductive Synthesis (IDS) jointly synthesizes implementation and proof for formally verified distributed systems, where SOTA agents (Codex/GPT-5.4, Claude Opus 4.6) succeed on only 2/7 tasks. Strong author list and significant capability gap addressed.

AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage that testing alone cannot provide. Distributed systems are a prime example: properties such as consistency between reads and writes must hold under every possible interleaving of events. Mechanized formal verification can guarantee such correctness, but typically demands months to years of expert effort. As evidence, even SOTA coding agents (Codex
AI AgentsFormal VerificationCode Generation
Research arXiv (Machine Learning) May 25

Test-Time Training Undermines Safety Guardrails

By Simone Antonelli, Sadegh Akhondzadeh, Aleksandar Bojchevski

75 score
AI Analysis

Demonstrates that Test-Time Training enables new jailbreak attacks with 95% Attack Success Rate over 10 trials under LoRA, transferring across model families. Important safety finding for adaptive inference.

Test-Time Training (TTT) is an emerging paradigm that enables models to adapt their parameters during inference, improving performance on tasks such as few-shot learning, retrieval-augmented generation, and complex reasoning. However, this dynamic adaptation introduces new vulnerabilities that adversaries can exploit to jailbreak models. We identify three threat models for TTT and demonstrate how attackers can leverage them to bypass safety filters. Our results show that TTT can significantly in
AI SafetyJailbreakingTest-Time Training
Research arXiv (Machine Learning) May 25

Decomposing and Measuring Evaluation Awareness

By Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Sahar Abdelnabi, Maksym Andriushchenko

75 score
AI Analysis

Decomposes evaluation awareness into environment recognizability and model propensity, operationalizing through 8 trigger factors and CoT monitoring across 9 frontier models and 4 benchmarks. Important framework for studying evaluation gaming.

Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results. Yet the field studies it without a shared foundation, conflating properties of the evaluation with properties of the model, and detection with behavioral response. We ground evaluation awareness in social psychology, decomposing it into an environment component (how recognizable the task is) and a model component that separates recognition from propensi
AI SafetyEvaluation AwarenessAlignment
Research arXiv (Machine Learning) May 25

Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization

By Alexander Long, Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Yan Zuo, Gil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi, Sameera Ramasinghe

72 score
AI Analysis

UPMs introduce time-varying invertible transforms at participant boundaries in distributed model training so that no participant ever holds extractable weights, while preserving network function. Tested on Qwen-2.5 and Llama-3.2 with negligible perplexity loss.

We consider a decentralized setup in which the participants collaboratively train and serve a large neural network, and where each participant only processes a subset of the model. In this setup, we explore the possibility of unmaterializable weights, where a full weight set is never available to any one participant. We introduce Unextractable Protocol Models (UPMs): a training and inference framework that leverages the sharded model setup to ensure model shards (i.e., subsets) held by participa
Distributed TrainingPrivacyModel Security
Research arXiv (Machine Learning) May 25

Approaching I/O-optimality for Approximate Attention

By P\'al Andr\'as Papp, Aleksandros Sobczyk, Anastasios Zouzias

70 score
AI Analysis

Presents I/O-efficient algorithms for approximate attention with almost-linear (rather than quadratic) I/O cost in sequence length n, improving on FlashAttention's complexity. Theoretically significant for scaling LLMs to long contexts.

We revisit the I/O complexity of attention in large language models. Given query-key-value matrices $Q,K,V\in\mathbb{R}^{n\times d}$, and a machine with fast memory size $M$, the goal is to compute the "attention matrix" $A=\text{softmax}(Q K ^{\top}/\sqrt{d}) V$ with the minimal number of data transfers between fast and slow memory. Existing methods in the literature, most notably FlashAttention and its variants, incur an I/O cost that depends quadratically on $n$, while a trivial lower bound o
EfficiencyAttentionTheory

Current evidence

Social Media

View category →

AI community chatter on 2026-05-24 centered on open-source friction, coding agent workflows, and renewed skepticism about the GenAI bubble.

85 score
AI Analysis

vLLM project bans contributor for 'AI slop' PR submitted as part of resume-building 'PR training' workflow; announces formal channel for important contributions and warns about AI-generated low-quality OSS contributions.

Thanks to the community report, we recently identified a PR t.co/QWboSmskkF that attempted to solve a non-existent issue and was submitted as part of a “PR training” workflow for resume building. The contributor involved has been banned from the vLLM community. This kind of low-signal contribution increases maintainer review overhead and creates unnecessary operational costs for open-source projects. As AI coding agents make generating large volumes of small PRs increasingly cheap, op
open sourceAI slopvLLMOSS sustainability
80 score
AI Analysis

Cherny's top tip for Claude Code: use auto mode (no permission prompts) to enable parallel multi-clauding

People often ask what my biggest tip is for getting the most out of Claude Code. These days my #1 tip is: use auto mode Auto mode means no more permission prompts. It is the key building block for multi-clauding: start a session, then while it runs, work on another session in parallel.
Claude Codeauto modeagentic workflowsproductivity
75 score
AI Analysis

Mollick predicts mass realization of how much online content is AI-generated.

As more people come to recognize the tells of AI, which mostly happens as you start to work with AI a lot, the scales are going to fall from their eyes and they are going to realize what some of us already see: how much of this site (and blog posts, articles, papers) are AI now.
AI contentAI detectiononline discourse