Is This AGI? Google’s Gemini 3 Deep Think Shatters Humanity’s Last Exam And Hits 84.6% On ARC-AGI-2 Performance Today
By Michal Sutter
Google's Gemini 3 Deep Think update achieves 84.6% on ARC-AGI-2, a benchmark considered a frontier test of general reasoning. The model uses extended test-time compute ('thinking longer') and internal verification to solve problems previously requiring human expert intervention. This represents a major leap toward AGI-class reasoning capabilities.