Daily AI intelligence

Daily AI Briefing — June 25, 2026

1570 current signals analyzed across AI news, research, social media, and open-source projects.

Daily synthesis

Executive Summary

Top Story

OpenAI and Broadcom unveiled Jalapeño, OpenAI's first custom ASIC designed from scratch over roughly nine months specifically for large language model inference at scale, marking a full-stack vertical-integration push to control inference economics for ChatGPT, Codex, and agentic workloads.

Key Developments

Safety & Regulation

Research Highlights

Looking Ahead

Watch whether OpenAI's custom silicon and the wave of multi-billion-dollar compute deals deliver durable cost advantages, or instead feed the intensifying debate over an AI infrastructure bubble.

Cross-category signals

Top Topics

Top Topic

OpenAI's Jalapeño Inference Chip

OpenAI and Broadcom announced Jalapeño, OpenAI's first custom ASIC designed from scratch over roughly nine months specifically for large language model inference at scale, with both Ars Technica and TechCrunch covering the launch. OpenAI's official account and president Greg Brockman drove massive social engagement (around 3.8M views), touting strong performance-per-watt and noting OpenAI's own models helped accelerate the chip's design to power ChatGPT, Codex, and agentic workloads. On Reddit, r/accelerate hyped the full-stack vertical-integration milestone while r/LocalLLaMA users questioned the performance-per-watt claims.
3 Social 2 News

Top Topic

AI Compute Buildout & Bubble Fears

Qualcomm agreed to acquire AI-software startup Modular for nearly $4 billion to expand from edge devices into data centers, while Blackstone committed $30 billion to AI data centers in Japan, per Wired and AI Business. Cerebras shares slid after its first post-IPO earnings on margin concerns. On social media, Gary Marcus amplified a Bloomberg quote calling the AI surge the largest stock-market bubble in history, underscoring growing skepticism about the scale of capital and leverage flowing into AI infrastructure.
3 News 1 Social

Top Topic

Agentic AI & Workplace Agents

Anthropic launched Claude Tag, embedding its AI in Slack and claiming it already writes 65 percent of the company's internal code, per The Decoder. On social media, Andrej Karpathy clarified his 'org-level harness' concept as a multiplayer agentic-work paradigm, Harrison Chase detailed a LangSmith 'sleep-time compute' pattern for agent memory, and Allie Miller described per-teammate Claude triage channels. Research reinforced the theme with Autodata's agentic data scientist, while Reddit's r/LocalLLaMA dug into Qwen-AgentWorld, a 3B-active MoE world model that simulates terminal, OS, SWE, and MCP environments for agents.
3 Social 1 News 1 Research

Top Topic

Chinese Open Models Pressure Labs

Snowflake's CEO reported that Zhipu AI's GLM-5.2 nearly matched Claude Opus 4.7 on a 103-task coding benchmark at roughly one-fifth the cost per output token, per The Decoder. On social media, Nathan Lambert praised GLM's wins while flagging its brittle, jagged behavior versus closed models and arguing it pressures frontier-lab margins, and swyx noted Zhipu's (Zai) upcoming IPO and GLM overtaking DeepSeek. Reddit's r/LocalLLaMA shared GLM5.2 optimization hacks boosting throughput from about 2.5 to over 50 tokens per second on a GH200 system.
2 Social 1 News

Top Topic

AI Safety & Over-Alignment

A rich cluster of safety research landed today: one paper found that thinking tokens do not reliably improve refusal behavior across frontier open-weight reasoning models, MATS-mentored work showed reward hacking generalizing toward broader misalignment when training Kimi K2.5 and GPT-OSS 120b, and the TSJ longitudinal framework exposed cognitive-developmental risks of AI companions for minors. A separate study demonstrated that editing just 125 Wikipedia articles can measurably shift LLM values, exposing a concrete data-poisoning vector. On Reddit, r/LocalLLaMA discussed the Swiss Federal Supreme Court evaluating Heretic to curb over-alignment, alongside debate over the Mythos/Fable export ban.
4 Research

Top Topic

AI Talent & IP Disputes

Senior Google AI researchers Jonas Adler and Alexander Pritzel are leaving for Anthropic, continuing a wave of departures, per TechCrunch, while a roughly $27 million Anthropic–OpenAI political proxy war over congressional candidate Alex Bores ended in a draw, per The Verge. On Reddit, r/LocalLLaMA debated Anthropic's public accusation that Alibaba ran a campaign using around 25,000 fraudulent accounts to illicitly distill Claude's capabilities, splitting the community on whether distillation is theft or fair game.
2 News

Current evidence

AI News

View category →

Compute infrastructure dominated the day. The move signals frontier labs vertically integrating to control inference economics.

Mistral shipped OCR 4 (claimed wins in 72% of blind tests), and Gradium launched speech-translation models it says beat gpt-realtime-translate.

News Ars Technica - All content Jun 24

OpenAI and Broadcom announce chip designed for LLM inference at scale

By Samuel Axon

80 score
AI Analysis

OpenAI and Broadcom announced Jalapeno, a custom ASIC designed specifically for large language model inference in data centers. Both companies frame it as the first generation of a long-term silicon roadmap aimed at running models at scale.

OpenAI, the company behind ChatGPT and Codex and the models those tools utilize, and Broadcom, an established silicon supplier, have announced a new chip called Jalapeño, designed specifically for large language model inference in data centers. The chip is intended to be deployed at large data centers, both companies claim this is just the first generation in a long-term project that will see chips refined over time.Read full article Comments
AI HardwareCompute InfrastructureOpenAI
News AI News & Artificial Intelligence | TechCrunch Jun 24

OpenAI unveils its first custom chip, built by Broadcom

By Russell Brandom

78 score
AI Analysis

TechCrunch reports OpenAI unveiled its first custom processor, Jalapeno, built with Broadcom and tailored for the specific demands of OpenAI inference systems. It marks OpenAI entry into bespoke AI silicon.

Named Jalapeño, the new processor was designed specifically for the unique needs of OpenAI's inference systems.
AI HardwareOpenAICompute Infrastructure
News Feed: Artificial Intelligence Latest Jun 24

Qualcomm Buys Buzzy Chip Startup Modular for Nearly $4 Billion

By Lauren Goode

70 score
AI Analysis

Qualcomm agreed to acquire AI chip software startup Modular for nearly 4 billion dollars, expanding from edge devices into data center infrastructure. Modular built one of the most prominent AI compiler and runtime stacks of the era.

Modular, one of the most promising chip software startups of the AI era, heads for a multibillion-dollar exit.
AI HardwareAcquisitionsCompute Infrastructure
News aibusiness Jun 24

Blackstone Commits $30B as Japan’s AI Battle Heats Up

By Graham Hope

66 score
AI Analysis

Blackstone committed 30 billion dollars to AI data center development in Japan as competition for AI infrastructure in the country intensifies. The plan joins a wave of multibillion-dollar data center deals.

The plan joins a host of other multi-billion-dollar AI data center deals.
Compute InfrastructureData CentersInvestment
58 score
AI Analysis

Continuing our coverage of the Claude Tag launch from yesterday, Anthropic launched Claude Tag, letting teams summon its AI in any Slack channel by tagging it and assigning tasks. The company says the tool already generates 65 percent of code on its product team.

Claude Tag lets teams bring Anthropic's AI into Slack by tagging @Claude in any channel and assigning it tasks. Internally, the tool already generates 65 percent of the code on Anthropic's product team, the company says. The article Claude Tag embeds Anthropic's AI in Slack, already writes 65 percent of internal code, company says appeared first on The Decoder.
Agentic AIAnthropicAI at Work

Current evidence

Research

View category →

Today's research centers on data quality, diffusion language models, and an unusually rich set of safety findings on deployed systems.

Data and training dynamics:

Architectures and scaling:

Safety and alignment:

Research arXiv (Artificial Intelligence) Jun 25

Internal Data Repetition Destroys Language Models

By Jessica Chudnovsky, Joshua Kazdan, Noam Levi, Rylan Schaeffer, Yegor Denisov-Blanch, Bo He, Mehmet Donmez, Sanmi Koyejo, David Donoho

74 score
AI Analysis

Revisits data repetition in the Chinchilla scaling era using Compute-Equivalent Gain/Loss, showing repetition damage is systematic and that repeating a moderate subset many times harms performance more than repeating a large subset few times. Important for data-constrained pretraining.

arXiv:2606.24998v1 Announce Type: cross Abstract: Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition. Earlier controlled studies predated Chinchilla-style scaling laws and could only measure the cost of repetition indirectly. We revisit repetition in the Chinchilla era, using a fitted no-repetition scaling law to report Compute-Equivalent Gain and Compute-Equivalent Loss. We show that under this modernized p
Scaling LawsPretrainingLanguage ModelsData
Research arXiv (Artificial Intelligence) Jun 25

Autodata: An agentic data scientist to create high quality synthetic data

By Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston

74 score
AI Analysis

Introduces Autodata, a method to train an agentic data scientist that creates high-quality synthetic training/eval data, with a meta-optimization that improves the agent itself, yielding gains over classical synthetic data methods. From a strong Meta FAIR-style author team.

arXiv:2606.25996v1 Announce Type: new Abstract: We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist agent, so that it learns to create even stronger data. We describe the overall formulation, and a specific practical implementation, Agentic Self-Instruct. We conduct experiments on computer science research tasks, legal reasoning tasks and reasoning
Agentic AISynthetic DataSelf-ImprovementLanguage Models
Research arXiv (Artificial Intelligence) Jun 25

Do Thinking Tokens Help with Safety?

By Narutatsu Ri, Abhishek Panigrahi, Sanjeev Arora

71 score
AI Analysis

Provides evidence that thinking tokens do not always improve safety: across frontier open-weight reasoning models, refusal/compliance outcomes are highly predictable from the first token's hidden representation, before any deliberation. Challenges the assumption that deliberation improves alignment.

arXiv:2606.25013v1 Announce Type: cross Abstract: Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts. It is also generally believed that this more "deliberative" mode should improve alignment and safety, by providing the model a safe space to consider whether its planned answer to a request violates its safety principles. We present evidence that this intuition is not always correct. Across frontier open-weight rea
AI SafetyReasoningAlignmentInterpretability
Research arXiv (Computation and Language) Jun 25

Real-Time Voice AI Hears but Does Not Listen

By Martijn Bartelds, Federico Bianchi, James Zou

70 score
AI Analysis

Evaluates four production real-time voice systems (GPT Realtime 2, Gemini 3.1 Flash Live, Qwen3.5 Omni Plus/Flash) on tasks where vocal delivery conveys meaning, finding they act on words and ignore distress, fear, or sarcasm. Notably the failure is in acting not perceiving, since systems can identify the emotion when asked directly.

arXiv:2606.26083v1 Announce Type: new Abstract: Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash-on tasks where the words and the delivery patterns both convey meaningful information. Across three consequential scenarios, all four systems act on the words rather than the voice. They end calls with crying callers who i
AI SafetySpeech ProcessingMultimodal ModelsAlignment
Research arXiv (Artificial Intelligence) Jun 25

Improved Large Language Diffusion Models

By Shen Nie, Qiyang Min, Shaoxuan Xu, Zihao Huang, Yuxuan Song, Yong Shan, Yankai Lin, Wayne Xin Zhao, Chongxuan Li, Ji-Rong Wen

70 score
AI Analysis

iLLaDA is an 8B masked diffusion language model trained fully from scratch with bidirectional attention, scaling pretraining to 12T tokens and improving substantially over the prior LLaDA on general, math, and code benchmarks. It demonstrates continued progress for diffusion-based alternatives to autoregressive LLMs.

arXiv:2606.25331v1 Announce Type: cross Abstract: Modern large language models are predominantly trained with autoregressive factorization and causal attention. We present \emph{iLLaDA}, an 8B masked diffusion language model trained from scratch with fully bidirectional attention. iLLaDA keeps the masked diffusion objective throughout pre-training and supervised fine-tuning (SFT), scaling pre-training to 12T tokens and fine-tuning on a 25B-token instruction corpus for 12 epochs. We further use
Language ModelsDiffusion ModelsPretraining

Current evidence

Social Media

View category →

OpenAI's custom silicon dominated the day. The company and president Greg Brockman unveiled Jalapeño, a from-scratch LLM-inference chip built with Broadcom over nine months, signaling full-stack vertical integration. The flagship announcement drew massive reach (3.8M views), with newsletters and commentators amplifying claims of strong inference performance powering ChatGPT, Codex, and agentic workloads.

85 score
AI Analysis

OpenAI announces Jalapeño, its first in-house AI chip built with Broadcom and purpose-built for LLM workloads powering ChatGPT, Codex, the API, and agentic products, framing it as expanding its full-stack platform.

We’ve designed and built our first AI chip: Jalapeño. Designed from the ground up by OpenAI and brought to production with @Broadcom, Jalapeño is purpose-built for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products. Chips are foundational to the AI economy. Building our own expands our full-stack platform from products to models to infrastructure, and will help us scale intelligence, serve more people, and expand access to AI.
AI hardwareOpenAIinference chipsAI infrastructure
78 score
AI Analysis

John Carmack reflects on early-career mistakes during Quake development, including over-ambition, pushing teams too hard at startup intensity, and poor founder stock arrangements.

There are a few things that I look back on as my mistakes in the early days. Quake was overly ambitious technically. We could have done all the great multiplayer and modding work inside a Doom++ engine, allowing the designers to work with a more stable base instead of rug-pulling everything out from underneath them a couple times. The follow up game could have then brought in full 6DOF environments and characters. I pushed everyone too hard. I didn’t appreciate how maturing companies need more
engineering leadershipcareer lessonsstartups
75 score
AI Analysis

Brockman introduces Jalapeño, an inference chip built from scratch over nine months and accelerated by OpenAI's own models, touting strong performance per watt.

Introducing Jalapeño — designed from scratch for LLM inference over nine months, accelerated by our models. Perf per watt looking incredible.
AI hardwareOpenAIinference chips
70 score
AI Analysis

Nathan Lambert shares a prerequisites lecture for his book covering language model basics, post-training masks, cross-entropy, KL divergence, and RL framing, built with GLM 5.2.

Another quick lecture -- I've been asked many times for prereq's to my book and what you should know, so built a little lecture (with GLM 5.2) to cover some more basics. Topics include: 00:00 Introduction & Course Prerequisites 01:37 Language Models Overview 02:47 The LM Head 04:29 Softmax & Log-Probabilities 06:13 Anatomy of an LM Training Example 06:37 Computing LLM Probabilities (+Phoebe the Dog) 09:52 Three Common Masks in Post-Training 11:03 A Small Decoding Review 12:14 Training an LM: C
AI educationML fundamentalspost-trainingreinforcement learning
62 score
AI Analysis

Continuing the discussion from yesterday, Karpathy explains his org-level harness is a genuinely different, enterprise-grade way of working, deeply integrated and multiplayer, where the system writes most code and everyone becomes a manager, distinct from Slack bots or OpenClaw.

@salomon_diei The basic idea is easy and v0 is a hackathon project. The product here is a lot closer to *it actually works*, for enterprise grade deployments, and after quite a bit of internal experimentation and iteration. It’s kind of hard to describe other than (per the post) it’s writing majority of code, it’s deeply integrated, multiplayer, and it starts to feel like everyone is a manager. So I understand it looks easy to dismiss on quick reading but it’s not some LLM Q&A with RAG over Slac
agentic toolsfuture of workdeveloper tooling