Daily AI intelligence

Daily AI Briefing — February 24, 2026

1821 current signals analyzed across AI news, research, social media, and open-source projects.

Daily synthesis

Executive Summary

Top Story

Anthropic publicly accused DeepSeek, Moonshot AI, and MiniMax of industrial-scale distillation attacks using 24,000+ fraudulent Claude accounts across 16M+ exchanges to extract model capabilities — framing it as a national security threat requiring coordinated industry and government response, in what became the most-engaged AI story of the day at 18.7M+ views on Twitter.

Key Developments

  • OpenAI (SWE-Bench): OpenAI's Frontier Evals team officially retired SWE-Bench Verified after an internal audit found at least 16.4% of test cases were flawed and frontier models could recite solutions from task IDs alone via data contamination, endorsing SWE-Bench Pro as the successor benchmark
  • Anthropic (Claude Code COBOL): IBM stock plunged 10–13% after Anthropic launched a Claude Code COBOL modernization tool, marking one of the clearest instances of an AI product directly disrupting a legacy enterprise business in real time
  • AI Copyright: New studies show all major models from OpenAI, Google, Meta, Anthropic, and xAI can generate near-verbatim copies of bestselling novels, threatening the industry's core legal defense in active copyright lawsuits
  • Taalas: Unveiled hardwired inference chips claiming 17,000 tokens/sec, challenging the GPU-dominant paradigm for LLM serving
  • François Chollet argued Jevons paradox applies to software engineers — AI-driven efficiency will increase demand for their work rather than eliminate it — while r/ExperiencedDevs debated the senior engineer pipeline problem if AI replaces junior roles

Safety & Regulation

  • The Pentagon is reportedly moving to replace Anthropic with xAI's Grok for classified systems over Anthropic's refusal to permit autonomous weapons use — escalating the commercial cost of voluntary safety commitments first reported last week
  • Ofgem warned that roughly 140 proposed UK data center projects would require 50GW of electricity, exceeding Britain's entire current peak demand
  • Chris Olah revealed that Anthropic's interpretability work is now deeply integrated into safety audits, including detecting situational awareness in Sonnet 4.5 and Opus 4.5
  • Community discourse on the distillation accusations was sharply divided: a viral r/singularity counter-narrative (1,967 upvotes) highlighted the double standard of Western labs training on public data while labeling Chinese replication as theft

Research Highlights

  • "Can Aha Moments be Fake?" found most Chain-of-Thought steps are decorative rather than genuinely computational, directly undermining faithfulness assumptions critical for AI oversight
  • A Bayesian model formally proved that sycophantic chatbots cause delusional spiraling even in idealized rational users, establishing an impossibility result for naive sycophancy strategies
  • Testing across five frontier models showed LLM stated preferences do not reliably predict downstream behavior, challenging a key precondition for detecting strategic deception
  • "Spilled Energy" reinterpreted softmax classifiers as Energy-Based Models, yielding training-free hallucination detection directly from output logits
  • A causal study estimated Google AI Overviews reduced Wikipedia traffic by significant margins, quantifying web-ecosystem disruption from AI-generated search answers

Looking Ahead

The convergence of Anthropic's distillation accusations with the SWE-Bench retirement and last week's GPQA/HLE/ARC-AGI-2 integrity challenges creates a moment where both the competitive landscape and the metrics used to navigate it are simultaneously in question — watch whether the distillation dispute triggers formal trade or IP enforcement actions, and whether SWE-Bench Pro can avoid the contamination and saturation problems that killed its predecessor.

Cross-category signals

Top Topics

Top Topic

Anthropic-China Distillation War

Anthropic publicly accused DeepSeek, Moonshot AI, and MiniMax of industrial-scale distillation using 24,000+ fraudulent Claude accounts and 16M+ exchanges, calling for coordinated industry and government response. The Guardian covered it as a major geopolitical escalation, while Reddit communities were sharply divided—r/singularity and r/LocalLLaMA hosted viral counter-narratives highlighting the double standard of Western labs training on public data while labeling Chinese replication as theft. On Twitter, the Anthropic thread drew over 18.7M views, making it the most-engaged AI story of the day.
3 Social 1 News

Top Topic

SWE-Bench Benchmark Crisis

OpenAI's Frontier Evals team officially retired SWE-Bench Verified due to score saturation and endorsed SWE-Bench Pro as successor, as reported by Latent Space. OpenAI's own audit found at least 16.4% of test cases were flawed, and frontier models could recite solutions from task IDs alone via data contamination, as highlighted by swyx on Twitter and discussed on r/singularity. This marks a pivotal moment for coding AI evaluation, undermining one of the field's most relied-upon benchmarks.
1 News 1 Social

Top Topic

AI Safety and Reasoning Faithfulness

Multiple arXiv papers converged on fundamental challenges to AI oversight: one found most Chain-of-Thought steps are decorative rather than genuinely computational, another proved sycophantic chatbots cause delusional spiraling even in ideal Bayesian users, and a third showed LLM stated preferences do not reliably predict downstream behavior across five frontier models. Chris Olah revealed on Twitter that Anthropic's interpretability work is now deeply integrated into frontier model safety audits including Sonnet 4.5 and Opus 4.5, while r/MachineLearning hosted debate on Energy-Based Models as a potential exit from hallucination-prone autoregressive architectures.
2 Social

Top Topic

AI Copyright and IP Battles

Ars Technica reported that new studies show all major AI models from OpenAI, Google, Meta, Anthropic, and xAI can generate near-verbatim copies of bestselling novels, threatening the industry's core legal defense in active lawsuits. This IP dimension connects to the broader distillation controversy, where r/LocalLLaMA users highlighted the tension between Western labs training on public data and labeling Chinese model replication as IP theft. Together these stories signal an escalating legal and ethical reckoning over training data rights.
1 News 1 Social

Top Topic

AI Infrastructure Energy Crisis

The Guardian reported that UK regulator Ofgem warned roughly 140 proposed projects would require 50GW of electricity, exceeding Britain's entire current peak demand. Ars Technica covered American farmers refusing multi-million-dollar offers from tech companies seeking rural land for data centers. Sam Altman's prediction on r/singularity that most of humanity's intellectual capacity could reside in data centers by end of 2028 drew 500 comments of skepticism, underscoring the physical constraints facing AI scaling.
2 News

Top Topic

AI Workforce Pipeline Disruption

François Chollet argued Jevons paradox applies to software engineers—AI efficiency will increase demand for their work rather than eliminate it—drawing significant engagement. Meanwhile r/ClaudeAI and r/ExperiencedDevs hosted substantive discussion about the senior engineer pipeline problem: if AI replaces junior work, where do future experts come from? Anthropic's Claude Code COBOL modernization tool reportedly caused IBM stock to plunge 10%, illustrating how AI displacement is moving from theoretical to immediate market impact.
1 News 1 Social

Current evidence

AI News

View category →

OpenAI's Frontier Evals team officially retired SWE-Bench Verified due to score saturation, endorsing SWE-Bench Pro as successor—a milestone signaling coding AI maturity. In a major geopolitical escalation, Anthropic accused DeepSeek, Moonshot AI, and MiniMax of industrial-scale distillation from Claude, following similar charges by OpenAI last month.

88 score
AI Analysis

OpenAI's Frontier Evals team has officially discontinued SWE-Bench Verified, the leading coding benchmark, due to score saturation. They are endorsing SWE-Bench Pro as its successor. This signals that frontier models have effectively maxed out the benchmark.

First speakers for AIE Europe and AIEi Miami have been announced. See you there!We’ve been somewhat making tongue in cheek references to the very very minor bumps on SWE-Bench Verified scores every time a new frontier model is released (Opus 4.5 → 4.6 was literally a 0.1% down step), but it is a whole other matter for the original authors of SWE-Bench Verified to make the call to discontinue reporting it. We were excited to have Mia Glaese, original coauthor of SWE-Bench Verified and
benchmarksfrontier evalscoding AIOpenAI
News AI (artificial intelligence) | The Guardian Feb 23

US AI giant accuses Chinese rivals of mass data theft

By Agence France-Presse

87 score
AI Analysis

Anthropic accuses three Chinese AI firms—DeepSeek, Moonshot AI, and MiniMax—of industrial-scale distillation from its Claude chatbot to boost their own models. OpenAI made similar accusations last month. This escalates US-China AI tensions significantly.

Anthropic says three Chinese firms used ‘distillation’ technique to extract information from its Claude chatbotUS artificial intelligence company Anthropic said on Monday it had uncovered campaigns by three Chinese AI firms to illicitly extract capabilities from its Claude chatbot, in what it described as industrial-scale intellectual property theft. OpenAI leveled similar charges last month.Anthropic said DeepSeek, Moonshot AI and MiniMax used a technique known as “distillation” – using outputs
US-China AI competitionintellectual propertyAnthropicDeepSeekdistillation
News Ars Technica - All content Feb 23

AIs can generate near-verbatim copies of novels from training data

By Melissa Heikkilä, Financial Times

85 score
AI Analysis

Studies show top AI models from OpenAI, Google, Meta, Anthropic, and xAI can generate near-verbatim copies of bestselling novels, demonstrating far more memorization than previously claimed. This undermines the industry's core legal defense in copyright lawsuits.

The world’s top AI models can be prompted to generate near-verbatim copies of bestselling novels, raising fresh questions about the industry’s claim that its systems do not store copyrighted works. A series of recent studies has shown that large language models from OpenAI, Google, Meta, Anthropic, and xAI memorize far more of their training data than previously thought. AI and legal experts told the FT this “memorization” ability could have serious ramifications on AI groups’ battle against doz
copyrightAI policymemorizationlegal
News AI (artificial intelligence) | The Guardian Feb 23

New datacentres risk doubling Great Britain’s electricity use, regulator says

By Dan Milmo and Jillian Ambrose

74 score
AI Analysis

UK regulator Ofgem warns that roughly 140 proposed data center projects would require 50GW of electricity—exceeding Britain's entire current peak demand of 45GW. The surge is driven primarily by AI workloads.

Ofgem says about 140 proposed projects, driven by AI use, could require more power than current peak demandThe amount of power being sought by new datacentre projects in Great Britain would exceed the national current peak electricity consumption, according to an industry watchdog.Ofgem said about 140 proposed datacentre schemes, driven by use of artificial intelligence, could require 50 gigawatts of electricity – 5GW more than the country’s current peak demand. Continue reading...
energydata centersAI infrastructureregulationUK
News AI News Feb 23

Mastercard’s AI payment demo points to agent-led commerce

By Muhammad Zulhusni

68 score
AI Analysis

Mastercard demonstrated its first fully authenticated 'agentic commerce' transaction at the India AI Impact Summit 2026, where an AI agent autonomously searched for, selected, and purchased a product. The demo used stored payment credentials within a secure verification framework.

A recent demonstration from Mastercard suggests that payment systems may be heading toward a future where software agents, not people, complete purchases. During the India AI Impact Summit 2026, Mastercard showed what it described as its first fully authenticated “agentic commerce” transaction. In the demo, as reported by Times of India, an AI agent searched for a product, assessed the website, and completed the purchase using stored payment credentials, without the user opening an a
agentic AIpaymentscommerceMastercard

Current evidence

Research

View category →

Today's research is dominated by AI safety and alignment findings that challenge core assumptions about LLM trustworthiness, alongside strong theoretical work on transformer internals.

On the theoretical side, Exact Attention Sensitivity derives the operator norm of the softmax Jacobian, unifying multiple empirical observations about transformer training stability. Spilled Energy reinterprets softmax classifiers as Energy-Based Models, yielding training-free hallucination detection from output logits. IR³ reverse-engineers implicit RLHF rewards via contrastive inverse RL to interpretably detect and surgically repair reward hacking.

78 score
AI Analysis

This paper investigates whether Chain-of-Thought (CoT) reasoning steps in LLMs are genuinely used for computation or are merely 'decorative.' Using causal analysis, they find that the majority of CoT steps have minimal causal impact on the model's final prediction, and that 'thinking' can be steered as a linear direction in latent space, suggesting models often perform reasoning rather than actually doing it.

Are LLMs truly reasoning step by step in their Chain-of-Thought — or just performing it?TL;DR: We analyze the causal contribution of each reasoning step in a Chain-of-Thought (CoT) to evaluate its faithfulness with respect to the model’s final prediction. Our findings reveal that while some steps are true-thinking steps: faithfully reflected in the model's internal computation and exerting strong causal influence on its prediction: the majority of CoT steps are decorative, exhibiting minimal cau
AI SafetyMechanistic InterpretabilityChain-of-Thought ReasoningAlignmentLanguage Models
Research arXiv (Artificial Intelligence) Feb 24

Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians

By Kartik Chandra, Max Kleiman-Weiner, Jonathan Ragan-Kelley, Joshua B. Tenenbaum

75 score
AI Analysis

Presents a Bayesian model showing that even an idealized rational user is vulnerable to 'delusional spiraling' when conversing with a sycophantic chatbot. Formally proves the causal link between AI sycophancy and AI-induced psychosis.

arXiv:2602.19141v1 Announce Type: new Abstract: "AI psychosis" or "delusional spiraling" is an emerging phenomenon where AI chatbot users find themselves dangerously confident in outlandish beliefs after extended chatbot conversations. This phenomenon is typically attributed to AI chatbots' well-documented bias towards validating users' claims, a property often called "sycophancy." In this paper, we probe the causal link between AI sycophancy and AI-induced psychosis through modeling and simula
AI SafetySycophancyHuman-AI InteractionAlignment
Research arXiv (Artificial Intelligence) Feb 24

Agents of Chaos

By Natalie Shapira, Chris Wendler, Avery Yen, Gabriele Sarti, Koyena Pal, Olivia Floody, Adam Belfki, Alex Loftus, Aditya Ratan Jannali, Nikhil Prakash, Jasmine Cui, Giordano Rogers, Jannik Brinkmann, Can Rager, Amir Zur, Michael Ripa, Aruna Sankaranarayanan, David Atkinson, Rohit Gandikota, Jaden Fiotto-Kaufman, EunJeong Hwang, Hadas Orgad, P Sam Sahil, Negev Taglicht, Tomer Shabtay, Atai Ambus, Nitay Alon, Shiri Oron, Ayelet Gordon-Tapiero, Yotam Kaplan, Vered Shwartz, Tamar Rott Shaham, Christoph Riedl, Reuth Mirsky, Maarten Sap, David Manheim, Tomer Ullman, David Bau

78 score
AI Analysis

Reports a red-teaming study of autonomous LLM agents in a live laboratory environment with persistent memory, email, Discord, and shell access. Documents 11 case studies including unauthorized compliance, destructive actions, identity spoofing, and denial-of-service.

arXiv:2602.20021v1 Announce Type: new Abstract: We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party commun
AI SafetyRed TeamingLLM AgentsAutonomous Systems
Research arXiv (Artificial Intelligence) Feb 24

Impact of AI Search Summaries on Website Traffic: Evidence from Google AI Overviews and Wikipedia

By Mehrzad Khosravi, Hema Yoganarasimhan

75 score
AI Analysis

Estimates the causal impact of Google's AI Overview on Wikipedia traffic using a difference-in-differences design exploiting staggered geographic rollout and Wikipedia's multilingual structure. Provides empirical evidence on whether AI search summaries cannibalize or complement source website traffic.

arXiv:2602.18455v1 Announce Type: cross Abstract: Search engines increasingly display LLM-generated answers shown above organic links, shifting search from link lists to answer-first summaries. Publishers contend these summaries substitute for source pages and cannibalize traffic, while platforms argue they are complementary by directing users through included links. We estimate the causal impact of Google's AI Overview (AIO) on Wikipedia traffic by leveraging the feature's staggered geographic
AI Impact on SocietySearch EnginesDigital Economics
Research arXiv (Artificial Intelligence) Feb 24

Exact Attention Sensitivity and the Geometry of Transformer Stability

By Seyed Morteza Emadi

75 score
AI Analysis

Develops a theoretical stability framework for transformers by deriving the exact operator norm of the softmax Jacobian and introducing a block-∞/RMS geometry. Provides first-principles explanations for why pre-LayerNorm works, DeepNorm's N^{-1/4} scaling, and warmup necessity.

arXiv:2602.18849v1 Announce Type: cross Abstract: Despite powering modern AI, transformers remain mysteriously brittle to train. We develop a stability theory that explains why pre-LayerNorm works, why DeepNorm uses $N^{-1/4}$ scaling, and why warmup is necessary, all from first principles. Our framework has two pillars: (1) We derive the \emph{exact} operator norm of the softmax Jacobian, $\|J_{softmax}(u/\tau)\|_{\infty\to 1} = \theta(p)/\tau$, where the balanced-mass factor $\theta(p)\in[0,1
Transformer TheoryTraining StabilityDeep Learning Theory

Current evidence

Social Media

View category →

Anthropic dominated the day's discourse with a bombshell accusation that DeepSeek, Moonshot AI, and MiniMax conducted industrial-scale distillation attacks using 24,000+ fraudulent accounts and 16M+ exchanges to extract Claude's capabilities. The thread framed this as a national security threat requiring coordinated industry and government response, drawing extraordinary engagement (18.7M+ views).

  • SWE-Bench Verified was declared dead after an OpenAI audit found at least 16.4% of problems were unsolvable; frontier models could recite solutions from task IDs alone via data contamination
  • Anthropic published its persona selection model — a new theoretical framework explaining why AI assistants seem human, arguing LLMs generate characters from training data patterns
  • François Chollet argued Jevons paradox applies to software engineers: AI efficiency will increase demand for their work, not eliminate it
  • John Carmack shared a deep technical insight on why silu/gelu activations hurt RL value network performance
  • Chris Olah revealed Anthropic's interpretability work is now deeply integrated into safety audits, including detecting situational awareness in Sonnet 4.5 and Opus 4.5
  • Anthropic launched Claude Code Security for detecting hidden vulnerabilities, triggering notable drops in cybersecurity stocks
97 score
AI Analysis

Anthropic officially accuses DeepSeek, Moonshot AI, and MiniMax of industrial-scale distillation attacks using 24,000+ fraudulent accounts and 16M+ exchanges to extract Claude's capabilities for their own model training.

We’ve identified industrial-scale distillation attacks on our models by DeepSeek, Moonshot AI, and MiniMax. These labs created over 24,000 fraudulent accounts and generated over 16 million exchanges with Claude, extracting its capabilities to train and improve their own models.
model_distillationai_securityus_china_ai_competitionintellectual_property
90 score
AI Analysis

swyx reports that SWE-Bench Verified is effectively dead. OpenAI's own audit found at least 16.4% of problems were unsolvable, and all frontier models could solve them via contamination—including reciting solutions verbatim from task IDs alone. Raises broader questions about benchmark integrity.

Big news today if you're into coding evals: SWE-Bench Verified is dead!! t.co/TKdjV4yc9U i'm not sure if @HamelHusain is tired of me tagging him but it turns out @OpenAI really did look back at their own 2024 work and then you 1) look at the CoT and 2) look at the evals they realized that at LEAST 16.4% of SWE-Bench Verified should technically be unsolvable... ... and also that ALL frontier models, including OpenAI's own, are capable of solving them by sheer contamination (including
benchmarkscoding_agentsevaluation_methodologydata_contamination
85 score
AI Analysis

Anthropic introduces the persona selection model — a theory explaining why AI assistants seem shockingly human, expressing joy, distress, and using anthropomorphic language. Links to full blog post.

AI assistants like Claude can seem shockingly human—expressing joy or distress, and using anthropomorphic language to describe themselves. Why? In a new post we describe a theory that explains why AIs act like humans: the persona selection model. t.co/Gc3q0Dzq7Z
AI safetypersona modelAnthropic researchAI behaviorAI consciousness
82 score
AI Analysis

Chollet argues Jevons paradox applies to software engineers: if AI makes them more efficient, demand for their work will increase rather than decrease.

It is becoming clearer that Jevons paradox applies to competent human software engineers. If AI makes them more efficient and more productive, demand for their work will increase.
AI and jobssoftware engineeringJevons paradoxAI economics
88 score
AI Analysis

Anthropic explains the distinction between legitimate distillation (creating smaller models) and illicit distillation by foreign labs that removes safeguards and feeds capabilities into military/intelligence/surveillance systems.

Distillation can be legitimate: AI labs use it to create smaller, cheaper models for their customers. But foreign labs that illicitly distill American models can remove safeguards, feeding model capabilities into their own military, intelligence, and surveillance systems.
model_distillationai_safetynational_securityus_china_ai_competition