Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
By Matthias Bastian
Continuing our coverage from [yesterday](/?date=2026-07-26&category=news#item-f3524e7a64d1), Anthropic's Claude Opus 5 has shattered ARC-AGI-3 benchmark records, scoring 30.2 percent and quadrupling previous highs set by GPT-5.6 Sol. Researchers highlighted spontaneous logical reflection behaviors previously unseen in language models.