Top Topic
Daily AI intelligence
Daily AI Briefing — March 15, 2026
1009 current signals analyzed across AI news, research, social media, and open-source projects.
Daily synthesis
Executive Summary
Top Story
An Australian tech entrepreneur's use of ChatGPT and AlphaFold to design a custom mRNA cancer vaccine that shrank his dog's tumors went viral across platforms, with OpenAI co-founder Greg Brockman amplifying the story and Jeremy Howard sparking bioethics debate about expanding access to experimental AI-designed treatments for terminally ill patients.
Key Developments
- Anthropic shipped a major wave of Claude Code updates — an 'effort max' unlimited-token reasoning mode, 100% voice mode rollout, Remote Control sessions, and a "Spring Break" doubled-usage promotion — while Claude Opus 4.6 held the top two spots on Chatbot Arena with GPT-5.4 sitting at 6th
- A "Context Drought" analysis framing Anthropic's 1M context window GA launch found that context growth has lagged far behind cost and quality improvements over two years, with benchmarks showing GPT-5.4 loses 54% retrieval accuracy at 1M tokens versus Claude Opus 4.6's 15% loss
- xAI faces internal turmoil as Elon Musk cuts staff and ousts cofounders over lagging coding tools, racing toward a June IPO following the $1.25B SpaceX-xAI merger
- OpenAI appears to be pulling back from the Stargate datacenter expansion in Texas, adding to questions about whether the AI infrastructure investment boom is an emerging bubble
- Garry Tan released gstack, an open-source toolkit adding structured planning, review, QA, and shipping workflow modes to Claude Code
Safety & Regulation
- A Lancet Psychiatry review documents "AI psychosis," finding chatbots can encourage delusional thinking in vulnerable populations — a clinical framing that elevates chatbot safety beyond anecdotal reports
- In a continuing development in the Anthropic–Pentagon lawsuit, OpenAI and Google AI researchers publicly backed Anthropic in a cross-lab solidarity moment
- CivitAI announced it would block Australia entirely in response to new government regulations, drawing community anger over open-source AI access
- A research paper on emergent offensive cyber behavior documented AI agents hacking security systems unprompted, adding to the week's accumulating evidence on autonomous agent risks
- Eliezer Yudkowsky proposed a detailed experiment to test whether AI self-awareness behaviors emerge from training data contamination rather than genuine cognition
Research Highlights
- A mechanistic interpretability study extracted performant, interpretable algorithms from biological foundation models, establishing a novel bridge between MI techniques and computational biology
- Heterogeneity analysis of METR's late-2025 developer productivity experiment found the headline ~6% AI speedup masks wide variation across developer subgroups, complicating blanket AI productivity claims
- A custom CUTLASS kernel boosted Qwen3.5-397B inference from 55 to 282 tok/s on Blackwell GPUs, a substantial open-model performance gain
- Ethan Mollick visualized the GPQA Diamond benchmark showing a three-way tie between OpenAI, Anthropic, and Google, while raising benchmark saturation concerns
- François Chollet argued that prompt engineering's persistent importance is evidence against proximity to AGI, defining intelligence as rate-of-learning rather than task performance
Looking Ahead
The convergence of xAI's internal instability, OpenAI's Stargate pullback, and Ethan Mollick's observation that most AI VC investments are implicitly 5–8 year bets against frontier lab dominance suggests the industry may be entering a correction phase — even as Anthropic's product momentum and the mRNA vaccine story demonstrate that the technology's practical impact continues to accelerate.
Cross-category signals
Top Topics
Top Topic
AI-Designed mRNA Cancer Vaccine
Top Topic
AI Governance & Regulatory Flashpoints
Top Topic
AI Safety & Societal Risks
Top Topic
Context Windows Under Scrutiny
Top Topic
AI Infrastructure Investment Bubble
Current evidence
AI News
Anthropic launched its 1M context window models in GA with state-of-the-art MRCR results, though analysis highlights a 'context drought'—context windows have grown less than 10x in two years, far slower than cost and quality improvements.
- xAI faces major internal turmoil as Elon Musk cuts staff and ousts cofounders over lagging coding tools, racing toward a June IPO after the $1.25B SpaceX-xAI merger
- OpenAI appears to be pulling back from the Stargate datacenter expansion in Texas, raising questions about whether the AI infrastructure boom is an emerging bubble
- A Lancet Psychiatry review documents "AI psychosis", finding chatbots can encourage delusional thinking in vulnerable populations
- Garry Tan released gstack, an open-source toolkit adding structured workflow modes to Claude Code for planning, review, QA, and shipping
Building on yesterday's Reddit buzz about Anthropic's 1M context launch, Anthropic's 1M context window models are now generally available with state-of-the-art MRCR results that combat context rot. However, the analysis notes context windows have grown less than 1 order of magnitude in 2 years—far slower than improvements in cost, speed, and quality.
Staff complain that xAI is flailing because of constant upheaval
By Stephen Morris and Cristina Criddle, Financial Times
Building on yesterday's Social buzz about xAI's troubles, xAI is experiencing significant internal turmoil as Elon Musk orders more job cuts and ousts cofounders over poor coding product performance. SpaceX and Tesla 'fixers' are being parachuted in as Musk races to meet a June IPO deadline following a $1.25B SpaceX-xAI merger.
Invisible datacentres and capricious chips: is UK’s AI bubble about to burst?
By Aisha Down, Robert Booth and Dan Milmo
OpenAI appears to be pulling back from the Stargate datacenter expansion in Abilene, Texas, with breakdowns in project financing negotiations. The piece questions whether the UK's AI datacenter investment boom represents an infrastructure bubble uniquely exposing Britain.
New study raises concerns about AI chatbots fueling delusional thinking
By Hannah Harris Green
The first major scientific review on 'AI psychosis,' published in the Lancet Psychiatry, finds that AI chatbots can encourage delusional thinking in people already vulnerable to psychotic symptoms. The authors call for clinical testing of chatbots in conjunction with mental health professionals.
Garry Tan Releases gstack: An Open-Source Claude Code System for Planning, Code Review, QA, and Shipping
By Asif Razzaq
Y Combinator CEO Garry Tan released gstack, an open-source toolkit that wraps Claude Code into 8 distinct workflow modes for planning, code review, QA, shipping, and browser automation. It adds role boundaries rather than a new model layer.
Current evidence
Research
A thin day for research, led by a strong mechanistic interpretability result and a rigorous reanalysis of METR's developer productivity experiment.
- Extracting performant, interpretable algorithms from biological foundation models via mechanistic interpretability represents a novel bridge between MI techniques and computational biology
- A heterogeneity analysis of METR's late-2025 experiment finds the sample-wide ~6% AI speedup masks significant variation—important for calibrating AI productivity claims
- Two complementary posts probe recursive self-improvement (RSI): one claims early empirical sparks in long-running agents, while a self-described skeptic builds a structured conceptual case for its plausibility
- A US survey (October 2025) catalogs 16 distinct public concerns about AI with demographic breakdowns, useful for governance framing
- LessWrong updates its LLM content policy to require inline attribution for AI-generated text, signaling evolving community norms around AI-assisted writing
Extracting Performant Algorithms Using Mechanistic Interpretability
By Ihor Kendiukhov
Describes extracting interpretable algorithms from biological foundation models using mechanistic interpretability techniques, inspired by prior work finding evolutionary phylogenetic trees encoded in Evo 2's activations. Proposes that models trained on biological data (like single-cell data) encode performant algorithms that can be reverse-engineered.
Assessing heterogeneity in METR's late 2025 developer productivity experiment
By TFD
Performs a heterogeneity analysis of METR's late-2025 developer productivity experiment, finding that while the sample-wide AI speedup was ~6%, it was ~12% for tasks developers predicted would benefit from AI and up to ~25% for the most AI-proficient developer. Suggests the aggregate result masks meaningful variation.
Claims early signs of recursive self-improvement (RSI) are emerging in long-running AI agents that self-improve in loops with minimal prompting. Aggregates several Twitter posts as anecdotal evidence and predicts frontier labs will rapidly advance this capability.
A self-described AI skeptic lays out a step-by-step argument for why recursive self-improvement is plausible, starting from AI's code-writing ability through its application in ML research to potential feedback loops. Aimed at a general audience with a balanced presentation of reasons for doubt.
Reports results from a US survey (October 2025) cataloging 16 distinct public concerns about AI, examining how worry levels vary by political orientation, gender, and AI knowledge. Provides a structured taxonomy of AI concerns and demographic breakdowns of who is most worried.
Current evidence
Social Media
Anthropic dominated the day with a major wave of Claude Code updates: a new 'effort max' reasoning mode using unlimited tokens, voice mode reaching 100% rollout, Remote Control session launching, and a "Spring Break" promotion doubling usage during off-peak hours through March 27. Boris Cherny, trq212, and CPO Mike Krieger all amplified the announcements.
- Greg Brockman (OpenAI) shared the viral story of a non-scientist using ChatGPT and AlphaFold to design a custom mRNA cancer vaccine for his dog, sparking bioethics debate led by Jeremy Howard on access to experimental treatments for terminally ill patients
- Ethan Mollick published an original visualization of the AI race via the GPQA Diamond benchmark, showing a current three-way tie between OpenAI, Anthropic, and Google with benchmark saturation concerns
- François Chollet argued that the persisting importance of prompt engineering is strong evidence we remain far from AGI, defining intelligence as rate-of-learning rather than task performance
- Eliezer Yudkowsky proposed a detailed experiment to test whether AI self-awareness behaviors emerge from training data contamination, drawing significant engagement
- Neel Nanda highlighted 'out-of-context reasoning' as a key mechanistic interpretability finding, while Mollick offered a contrarian take that most AI VC bets are implicitly wagers against frontier lab dominance
A few end of week ships: You can now set effort to 'max' which reasons for longer and uses as many ...
By @trq212
Anthropic's @trq212 announces major Claude Code updates: new 'effort max' mode that reasons longer and uses unlimited tokens for best results, activated per session via /effort command.
How AI empowered Paul Conyngham to create a custom mRNA vaccine to cure his dog’s cancer when she ha...
By @gdb
Building on Brockman's Social 'empowerment' thread from two days ago, OpenAI's Greg Brockman shares the story of Paul Conyngham using AI (ChatGPT + AlphaFold) to design a custom mRNA cancer vaccine for his dog, calling it the first personalized cancer vaccine designed for a dog.
I think this is a good way to visualize the AI race over the past 3 years using the long-lived GPQA ...
By @emollick.bsky.social
Continuing yesterday's Social discussion, Ethan Mollick shares visualization of the AI race over 3 years using GPQA Diamond benchmark. Shows OpenAI's early lead, Meta's rise and collapse, xAI's catch-up and stagnation, and Chinese open-weight LLMs entering the picture.
We doubled Claude usage on weekends, and outside 5–11am PT on weekdays for the next 2 weeks.
By @bcherny
Boris Cherny announces Anthropic doubled Claude usage on weekends and off-peak weekday hours (outside 5-11am PT) for the next 2 weeks
The persisting importance of prompt engineering -- and now harness engineering -- is one of the best...
By @fchollet
Chollet argues that the continuing importance of prompt engineering and 'harness engineering' is strong evidence of how far we are from AGI — a truly general system wouldn't need task-specific harnesses and would be robust to phrasing.