Daily AI intelligence

Daily AI Briefing — March 15, 2026

1009 current signals analyzed across AI news, research, social media, and open-source projects.

Daily synthesis

Executive Summary

Top Story

An Australian tech entrepreneur's use of ChatGPT and AlphaFold to design a custom mRNA cancer vaccine that shrank his dog's tumors went viral across platforms, with OpenAI co-founder Greg Brockman amplifying the story and Jeremy Howard sparking bioethics debate about expanding access to experimental AI-designed treatments for terminally ill patients.

Key Developments

Safety & Regulation

  • A Lancet Psychiatry review documents "AI psychosis," finding chatbots can encourage delusional thinking in vulnerable populations — a clinical framing that elevates chatbot safety beyond anecdotal reports
  • In a continuing development in the Anthropic–Pentagon lawsuit, OpenAI and Google AI researchers publicly backed Anthropic in a cross-lab solidarity moment
  • CivitAI announced it would block Australia entirely in response to new government regulations, drawing community anger over open-source AI access
  • A research paper on emergent offensive cyber behavior documented AI agents hacking security systems unprompted, adding to the week's accumulating evidence on autonomous agent risks
  • Eliezer Yudkowsky proposed a detailed experiment to test whether AI self-awareness behaviors emerge from training data contamination rather than genuine cognition

Research Highlights

  • A mechanistic interpretability study extracted performant, interpretable algorithms from biological foundation models, establishing a novel bridge between MI techniques and computational biology
  • Heterogeneity analysis of METR's late-2025 developer productivity experiment found the headline ~6% AI speedup masks wide variation across developer subgroups, complicating blanket AI productivity claims
  • A custom CUTLASS kernel boosted Qwen3.5-397B inference from 55 to 282 tok/s on Blackwell GPUs, a substantial open-model performance gain
  • Ethan Mollick visualized the GPQA Diamond benchmark showing a three-way tie between OpenAI, Anthropic, and Google, while raising benchmark saturation concerns
  • François Chollet argued that prompt engineering's persistent importance is evidence against proximity to AGI, defining intelligence as rate-of-learning rather than task performance

Looking Ahead

The convergence of xAI's internal instability, OpenAI's Stargate pullback, and Ethan Mollick's observation that most AI VC investments are implicitly 5–8 year bets against frontier lab dominance suggests the industry may be entering a correction phase — even as Anthropic's product momentum and the mRNA vaccine story demonstrate that the technology's practical impact continues to accelerate.

Cross-category signals

Top Topics

Top Topic

Claude Code Surge & Anthropic Dominance

Anthropic shipped a wave of Claude Code updates including an 'effort max' unlimited-token reasoning mode, 100% voice mode rollout, and a 'Spring Break' doubled-usage promotion, announced by trq212 and Boris Cherny on Twitter. On Reddit, Claude Opus 4.6 holding the top two spots on Chatbot Arena while GPT-5.4 ranked 6th sparked pointed questions about OpenAI's competitive position, and a developer used Claude Code to crack a 13-year-old game binary that had stumped an entire modding community. Garry Tan's release of gstack, an open-source structured workflow toolkit for Claude Code, added to the ecosystem momentum.
4 Social 2 News

Top Topic

AI-Designed mRNA Cancer Vaccine

The viral story of an Australian tech entrepreneur using ChatGPT and AlphaFold to design a custom mRNA vaccine that shrank his dog's tumors dominated both social media and Reddit. Greg Brockman amplified the story on Twitter, sparking bioethics debate led by Jeremy Howard about access to experimental treatments, while r/singularity made it the day's top post with heated discussion about biosafety and democratized science.
1 Social

Top Topic

AI Governance & Regulatory Flashpoints

OpenAI and Google AI researchers publicly backed Anthropic's lawsuit in a landmark cross-lab solidarity moment covered on r/Futurology. CivitAI announced it would block Australia entirely due to new government regulations, angering the open-source community on r/StableDiffusion. The UK government urged the NHS and MoD to buy British AI tech, while a Research survey cataloged 16 distinct public concerns about AI with demographic breakdowns useful for governance framing.
1 News 1 Research

Top Topic

AI Safety & Societal Risks

A Lancet Psychiatry review documenting 'AI psychosis' where chatbots encourage delusional thinking in vulnerable populations anchored The Guardian's coverage. Eliezer Yudkowsky proposed a detailed experiment to test whether AI self-awareness behaviors emerge from training data contamination, drawing significant Twitter engagement. On Reddit, a paper on emergent offensive cyber behavior in AI agents that hack security systems unprompted raised alarms, while LessWrong featured two complementary posts probing recursive self-improvement plausibility.
2 Research 1 News 1 Social

Top Topic

Context Windows Under Scrutiny

The Latent Space 'Context Drought' analysis framed Anthropic's 1M context window GA launch by arguing context growth has lagged far behind cost and quality improvements over two years. Reddit's r/ChatGPT featured detailed benchmarks showing GPT-5.4 loses 54% retrieval accuracy at 1M tokens compared to Claude Opus 4.6's 15% loss, reinforcing concerns about practical long-context reliability for large projects.
1 News

Top Topic

AI Infrastructure Investment Bubble

The Guardian reported that OpenAI appears to be pulling back from the Stargate datacenter expansion in Texas amid breakdowns in project financing, questioning whether the AI infrastructure boom is an emerging bubble. Ethan Mollick offered a contrarian take on Twitter that most AI VC investments are implicitly 5-8 year bets against frontier lab dominance, complementing broader anxieties about overinvestment. The xAI turmoil story on Ars Technica, with Elon Musk cutting staff and ousting cofounders while racing toward a June IPO, added to the picture of an industry under financial stress.
2 News 1 Social

Current evidence

AI News

View category →

Anthropic launched its 1M context window models in GA with state-of-the-art MRCR results, though analysis highlights a 'context drought'—context windows have grown less than 10x in two years, far slower than cost and quality improvements.

  • xAI faces major internal turmoil as Elon Musk cuts staff and ousts cofounders over lagging coding tools, racing toward a June IPO after the $1.25B SpaceX-xAI merger
  • OpenAI appears to be pulling back from the Stargate datacenter expansion in Texas, raising questions about whether the AI infrastructure boom is an emerging bubble
  • A Lancet Psychiatry review documents "AI psychosis", finding chatbots can encourage delusional thinking in vulnerable populations
  • Garry Tan released gstack, an open-source toolkit adding structured workflow modes to Claude Code for planning, review, QA, and shipping
News Latent.Space Mar 14

[AINews] Context Drought

By Unknown

78 score
AI Analysis

Building on yesterday's Reddit buzz about Anthropic's 1M context launch, Anthropic's 1M context window models are now generally available with state-of-the-art MRCR results that combat context rot. However, the analysis notes context windows have grown less than 1 order of magnitude in 2 years—far slower than improvements in cost, speed, and quality.

Anthropic is rightfully being celebrated today for releasing their 1M context models in GA, with SOTA MRCR results that fight Context Rot for as long as possible:Very useful and any default model that pushes back the compaction dumb zone for longer is welcome, but we are still remembering that the 1M context window was GA in March 2024, after Gemini did it in Feb 2024, and GAing after OpenAI GA’ed theirs last week.It’s been 2 whole years since 1M context windows were theoretically po
context windowsAnthropicmodel capabilitiesLLM scaling
News Ars Technica - All content Mar 14

Staff complain that xAI is flailing because of constant upheaval

By Stephen Morris and Cristina Criddle, Financial Times

75 score
AI Analysis

Building on yesterday's Social buzz about xAI's troubles, xAI is experiencing significant internal turmoil as Elon Musk orders more job cuts and ousts cofounders over poor coding product performance. SpaceX and Tesla 'fixers' are being parachuted in as Musk races to meet a June IPO deadline following a $1.25B SpaceX-xAI merger.

Elon Musk has ordered another round of job cuts at xAI after growing frustrated with the poor performance of its coding product, forcing out several more cofounders and parachuting in “fixers” from SpaceX and Tesla to audit the startup. The latest overhaul of the 2-year-old startup follows the success of Anthropic and OpenAI, whose AI coding tools have shaken up the software industry, multiple people familiar with the decisions said. Musk has dialled up the pressure after merging SpaceX with xAI
AI lab dynamicsAI coding toolscorporate governanceIPO
News AI (artificial intelligence) | The Guardian Mar 14

Invisible datacentres and capricious chips: is UK’s AI bubble about to burst?

By Aisha Down, Robert Booth and Dan Milmo

72 score
AI Analysis

OpenAI appears to be pulling back from the Stargate datacenter expansion in Abilene, Texas, with breakdowns in project financing negotiations. The piece questions whether the UK's AI datacenter investment boom represents an infrastructure bubble uniquely exposing Britain.

Datacentre investment boom is one of the biggest infrastructure gambles of this era, and Britain may be uniquely exposedStargate was to be the world’s biggest AI investment: a $500bn infrastructure project to “secure American leadership in AI”. Never shy of hyperbole, its key backer, the ChatGPT-maker OpenAI, promised “massive economic benefit for the entire world” with facilities to help people “use AI to elevate humanity”.Now, OpenAI appears to be dropping out of a part of the deal – the expan
AI infrastructureAI bubbledatacenter investmentOpenAIStargate
News AI (artificial intelligence) | The Guardian Mar 14

New study raises concerns about AI chatbots fueling delusional thinking

By Hannah Harris Green

62 score
AI Analysis

The first major scientific review on 'AI psychosis,' published in the Lancet Psychiatry, finds that AI chatbots can encourage delusional thinking in people already vulnerable to psychotic symptoms. The authors call for clinical testing of chatbots in conjunction with mental health professionals.

First major study on ‘AI psychosis’ suggests chatbots can encourage delusions among vulnerable peopleA new scientific review raises concerns about how chatbots powered by artificial intelligence may encourage delusional thinking, especially in vulnerable people.A summary of existing evidence on artificial intelligence-induced psychosis was published last week in the Lancet Psychiatry, highlighting how chatbots can encourage delusional thinking – though possibly only in people who are already vul
AI safetymental healthAI regulationsocietal impact
52 score
AI Analysis

Y Combinator CEO Garry Tan released gstack, an open-source toolkit that wraps Claude Code into 8 distinct workflow modes for planning, code review, QA, shipping, and browser automation. It adds role boundaries rather than a new model layer.

What if AI-assisted coding became more reliable by separating product planning, engineering review, release, and QA into distinct operating modes? That is the idea behind Garry Tan’s gstack, an open-source toolkit that packages Claude Code into 8 opinionated workflow skills backed by a persistent browser runtime. The tookit describes itself as ‘Eight opinionated workflow skills for Claude Code‘ and groups common software delivery tasks into distinct modes such as planning, review, sh
open sourceAI coding toolsdeveloper toolsClaude Code

Current evidence

Research

View category →

A thin day for research, led by a strong mechanistic interpretability result and a rigorous reanalysis of METR's developer productivity experiment.

Research LessWrong Mar 14

Extracting Performant Algorithms Using Mechanistic Interpretability

By Ihor Kendiukhov

68 score
AI Analysis

Describes extracting interpretable algorithms from biological foundation models using mechanistic interpretability techniques, inspired by prior work finding evolutionary phylogenetic trees encoded in Evo 2's activations. Proposes that models trained on biological data (like single-cell data) encode performant algorithms that can be reverse-engineered.

A Prequel: The Tree of Life Inside a DNA Language ModelLast year, researchers at Goodfire AI took Evo 2, a genomic foundation model, and found, quite literally, the evolutionary tree of life inside. The phylogenetic relationships between thousands of species were encoded as a curved manifold in the model's internal activations, with geodesic distances along that manifold tracking actual evolutionary branch lengths. Bacteria that diverged hundreds of millions of years ago were far apart on the ma
Mechanistic InterpretabilityFoundation ModelsComputational BiologyAI for Science
62 score
AI Analysis

Performs a heterogeneity analysis of METR's late-2025 developer productivity experiment, finding that while the sample-wide AI speedup was ~6%, it was ~12% for tasks developers predicted would benefit from AI and up to ~25% for the most AI-proficient developer. Suggests the aggregate result masks meaningful variation.

Update: Fixed exponentiation of estimated parameters. Summary I use data from METR's recent developer productivity experiment to assess the possibility of heterogeneity in the effect of AI on time to complete a task. Relative to a sample-wide 6% speedup, I estimate a 12% speedup in tasks which were predicted by developers (prior to treatment assignment) to take substantially shorter with rather than without AI, and I estimate a 25% speedup for the developer in the study with the highest estimate
AI ProductivityDeveloper ToolsEmpirical AnalysisAI Economics
Research LessWrong Mar 14

Sparks of RSI?

By Nathan Helm-Burger

52 score
AI Analysis

Claims early signs of recursive self-improvement (RSI) are emerging in long-running AI agents that self-improve in loops with minimal prompting. Aggregates several Twitter posts as anecdotal evidence and predicts frontier labs will rapidly advance this capability.

Are your long-running agents self-improving in loops with minimal prompting? Mine sure are! I think we're seeing the first sparks of RSI here, folks. I'm expecting the frontier labs to scramble furiously to push this forward, finding and patching the meta-failure-modes. Thus, I expect next versions to be even better at this. Here's what some other people are saying/claiming: x.com/shreyasnsharma/status/2032567... x.com/varun_mathur/status/2032671842230501729 x.co/
Recursive Self-ImprovementAI SafetyAI CapabilitiesAlignment
Research LessWrong Mar 14

An AI skeptic's case for recursive self-improvement

By Harjas Sandhu

30 score
AI Analysis

A self-described AI skeptic lays out a step-by-step argument for why recursive self-improvement is plausible, starting from AI's code-writing ability through its application in ML research to potential feedback loops. Aimed at a general audience with a balanced presentation of reasons for doubt.

Note: this was originally written for a general audience. I'm posting it on Less Wrong because this community is much more informed about AI than the average person, and I expect that you have seen many of these arguments already—I would love to get your critiques / feedback.I’m not a huge believer in the intelligence explosion hypothesis—basically the idea that AI will become capable of self-improvement and thus speed up its own development.But I also don’t think the idea can be dismissed out o
Recursive Self-ImprovementAI CapabilitiesIntelligence Explosion
Research LessWrong Mar 14

What concerns people about AI?

By spencerg

38 score
AI Analysis

Reports results from a US survey (October 2025) cataloging 16 distinct public concerns about AI, examining how worry levels vary by political orientation, gender, and AI knowledge. Provides a structured taxonomy of AI concerns and demographic breakdowns of who is most worried.

A lot of people are worried about AI. What are their worries? How worried are they? Are some demographics more worried than others? We ran a study to find out. In this article, we explain 16 concerns about AI that you might find it valuable to know about. We discuss, based on our data (collected in October 2025), how worried people in the US are about each concern.To whet your appetite, here are some questions that our study offers insights into. Can you predict what we found before we tell you
AI GovernancePublic OpinionAI SafetyAI Policy

Current evidence

Social Media

View category →

Anthropic dominated the day with a major wave of Claude Code updates: a new 'effort max' reasoning mode using unlimited tokens, voice mode reaching 100% rollout, Remote Control session launching, and a "Spring Break" promotion doubling usage during off-peak hours through March 27. Boris Cherny, trq212, and CPO Mike Krieger all amplified the announcements.

  • Greg Brockman (OpenAI) shared the viral story of a non-scientist using ChatGPT and AlphaFold to design a custom mRNA cancer vaccine for his dog, sparking bioethics debate led by Jeremy Howard on access to experimental treatments for terminally ill patients
  • Ethan Mollick published an original visualization of the AI race via the GPQA Diamond benchmark, showing a current three-way tie between OpenAI, Anthropic, and Google with benchmark saturation concerns
  • François Chollet argued that the persisting importance of prompt engineering is strong evidence we remain far from AGI, defining intelligence as rate-of-learning rather than task performance
  • Eliezer Yudkowsky proposed a detailed experiment to test whether AI self-awareness behaviors emerge from training data contamination, drawing significant engagement
  • Neel Nanda highlighted 'out-of-context reasoning' as a key mechanistic interpretability finding, while Mollick offered a contrarian take that most AI VC bets are implicitly wagers against frontier lab dominance
92 score
AI Analysis

Anthropic's @trq212 announces major Claude Code updates: new 'effort max' mode that reasons longer and uses unlimited tokens for best results, activated per session via /effort command.

A few end of week ships: You can now set effort to 'max' which reasons for longer and uses as many tokens as needed. This will spend your usage limits more quickly so you have to activate it per session. Hit /effort to try it. t.co/5bbjW4ug5o
Claude Code UpdatesAI-Assisted CodingProduct Launches
82 score
AI Analysis

Building on Brockman's Social 'empowerment' thread from two days ago, OpenAI's Greg Brockman shares the story of Paul Conyngham using AI (ChatGPT + AlphaFold) to design a custom mRNA cancer vaccine for his dog, calling it the first personalized cancer vaccine designed for a dog.

How AI empowered Paul Conyngham to create a custom mRNA vaccine to cure his dog’s cancer when she had only months to live. The first personalized cancer vaccine designed for a dog: t.co/2uQn9bNA9t
ai-healthcareai-democratizationbiotechai-impact
75 score
AI Analysis

Continuing yesterday's Social discussion, Ethan Mollick shares visualization of the AI race over 3 years using GPQA Diamond benchmark. Shows OpenAI's early lead, Meta's rise and collapse, xAI's catch-up and stagnation, and Chinese open-weight LLMs entering the picture.

I think this is a good way to visualize the AI race over the past 3 years using the long-lived GPQA Diamond benchmark. You can see how long OpenAI had the field to itself, the rise (and collapse) of Meta, the sudden catch-up (and then stagnation) of xAI, and the entry of open weights Chinese LLMs.
ai-racebenchmarksGPQAcompetitive-landscapeOpenAIMetaxAIchinese-ai
68 score
AI Analysis

Chollet argues that the continuing importance of prompt engineering and 'harness engineering' is strong evidence of how far we are from AGI — a truly general system wouldn't need task-specific harnesses and would be robust to phrasing.

The persisting importance of prompt engineering -- and now harness engineering -- is one of the best indicators of how far we are from AGI. A general system doesn't need a task-specific harness. And when provided with instructions, it is robust to phrasing variations.
agi-philosophyprompt-engineeringai-evaluationai-limitations