Category intelligence

Social Media Briefing — February 15, 2026

399 current items analyzed and ranked.

Executive synthesis

Social Media Summary

The AI community buzzed around OpenAI's "First Proof" benchmark, where models solved 6 of 10 novel unpublished math problems. Greg Brockman announced the results, and Sam Altman called AI producing genuinely new knowledge a significant milestone—while urging caution.

Key Themes

AI for Mathematics / First Proof Benchmark · 10DeepSeek v4 and Open Source AI Inflection · 6GPT-5.3-Codex & AI-Assisted Coding Revolution · 7Future of Software Engineering with AI · 8AI Reasoning vs Pattern Matching · 7AI Safety & Existential Risk · 10Human-AI Interaction Risks · 1AI IP & Data Rights · 1Chinese AI Market Dynamics · 2Claude Code Features and Development · 8

Primary evidence

Top Ranked Signals

92 score
AI Analysis

Building on yesterday's Reddit discussion about the First Proof results, Greg Brockman announces that OpenAI is benchmarking models on novel frontier research via 'First Proof' - their model found likely correct solutions to at least 6 of 10 unpublished math research problems in a week.

we are now benchmarking our models on novel frontier research, via t.co/2XmndVes5F. of 10 math research problems which research mathematicians have solved but never published the solutions to, in a week, our model discovered likely correct solutions to at least 6 of them.
AI for MathematicsFirst ProofOpenAI BreakthroughAI MilestonesBenchmarks
88 score
AI Analysis

Building on yesterday's Reddit discussion of the First Proof results, Sam Altman celebrates the progression from grade-school math to research-level math problems, calling this perhaps the most important eval now. Predicts dismissive reactions.

We went from AI systems that struggled to do grade school math to AI systems that can solve research-level math problems in just a few years. I agree with Jakub this is perhaps the most important eval now. I am also pretty sure the main reaction will be "it's not that hard" :)
AI for MathematicsAI MilestonesOpenAI StrategyBenchmarks
88 score
AI Analysis

Swyx announces a major shift in his stance on open source AI. He's been cynical for 3 years, notes Kimi K2.5 didn't beat GPT 5.2. Predicts DeepSeek v4 next week will be the turning point, mentions Chinese labs leaking information, other 'Tigers' lining up, and references 'Whalefall' as the stage being set.

i've been cynical on open source ai for the last 3 years, and it's not been a popular view. people want to hear that open source is catching up, that some underdog team found this One Weird Trick to outperform gpt5. Kimi K2.5 didnt even beat GPT 5.2 in the end. @DeepSeek_ai v4 next week is probably the moment I really change my stance for the first time. Hearing that the Chinese labs leak like a sieve (do you know which culture loves gossip more than Americans? that's right) and all the other T
open-source-aideepseek-v4chinese-ai-labsai-competitionmodel-releases
85 score
AI Analysis

Continuing from yesterday's Social announcement of GPT-5.2's physics result, Sam Altman says the ability to produce genuinely new knowledge is a significant milestone, urging both excitement and caution, while acknowledging the results aren't earth-shattering.

These are obviously not earth-shattering results, but the ability to produce genuinely new knowledge, however small, is a significant milestone and I hope we all take it seriously, with excitement and caution.
AI for MathematicsAI MilestonesOpenAI StrategyAI for Science
82 score
AI Analysis

Gary Marcus makes an urgent case for federal legislation banning AI impersonation of humans, citing Daniel Dennett's 'counterfeit people' essay, recent deepfake scams, and growing sophistication of voice/video synthesis tools.

We URGENTLY need a federal law forbidding AI from impersonating humans The night before I testified in the US Senate in May, 2023, the late philosopher Daniel Dennett sent me a manuscript that he called “counterfeit people”. It was published a few days later in The Atlantic. It started like this “MONEY HAS EXISTED for several thousand years, and from the outset counterfeiting was recognized to be a very serious crime, one that in many cases calls for capital punishment because it undermines
AI RegulationDeepfakesAI SafetyAI PolicyAI Ethics
82 score
AI Analysis

Boris Cherny (Anthropic) argues that engineering is changing, not dying: someone still needs to prompt Claudes, talk to customers, coordinate with teams, and decide what to build. Great engineers are more important than ever.

@big_duca Someone has to prompt the Claudes, talk to customers, coordinate with other teams, decide what to build next. Engineering is changing and great engineers are more important than ever.
future-of-engineeringai-coding-toolsanthropicclaude-codesoftware-engineering
Social Twitter Feb 14

how did we ever write all that code by hand

By @gdb

80 score
AI Analysis

Greg Brockman muses 'how did we ever write all that code by hand' - reflecting on the transformative nature of AI coding tools.

how did we ever write all that code by hand
AI-Assisted CodingGPT-5.3-Codex CapabilitiesSoftware Engineering Transformation
78 score
AI Analysis

Continuing Chollet's Social thread from yesterday on ARC benchmarks, Chollet reports that frontier models are overfitting to ARC's original encoding format due to extensive benchmark targeting. Performance remains tied to familiar input distributions.

Interesting finding on frontier model performance on ARC -- due to extensive direct targeting of the benchmark, models are overfitting to the original ARC encoding format. Frontier model performance remains largely tied to a familiar input distribution.
ARC BenchmarkBenchmark ValidityAI Reasoning vs Pattern MatchingOverfitting
76 score
AI Analysis

Following yesterday's Reddit discussion of the First Proof results, Emollick outlines the predictable trajectory of AI capabilities: from 'AI can't do X' to 'of course AI does X', mapping this pattern onto AI for novel science - over-enthusiastic claims, then human-AI collaboration, then increasing AI autonomy.

The transition from “AI can’t do novel science” to “of course AI does novel science” will be like every other similar AI transition. First the over-enthusiastic claims, then smart people use AI to help them, then AI starts to do more of the work, then minor discoveries, & then…
AI for ScienceAI Capability TrajectoryAI Milestones
75 score
AI Analysis

Following yesterday's Reddit discussion of the First Proof results, Emollick warns that as AI advances in math proofs, verification becomes a bottleneck that only a tiny number of humans can perform. Calls for solutions like multi-AI verification.

One of the things to watch out for as AI advances is the verifiably becomes something only a vanishingly small number of people can do (below is a mathematician on AI proofs) We need to start thinking harder about that problem (multiple AIs working together? something else?)
AI for MathematicsAI Verification ProblemAI Safety
Social Twitter Feb 14

GPT-5.3-Codex for UI design:

By @gdb

72 score
AI Analysis

Greg Brockman (OpenAI President) showcases GPT-5.3-Codex being used for UI design tasks.

GPT-5.3-Codex for UI design:
GPT-5.3-Codex CapabilitiesAI-Assisted DesignOpenAI Product
72 score
AI Analysis

Burkov defends DeepSeek (or another entity) against OpenAI's complaints about model distillation, arguing OpenAI trained on the entire internet without paying creators. Calls Altman hypocritical.

So what? OpenAI "distilled" the entire internet without paying the authors a penny to train its models. (I know. I wasn't paid.) altman, stop being a pussy. Stealing the stolen from a thief isn't theft. t.co/h1um0wSVyr
AI IP/CopyrightOpenAI CriticismAI EthicsData Rights