ARC-AGI-3 is out now! We've designed the benchmark to evaluate agentic intelligence via interactive ...
By @fchollet
François Chollet's main ARC-AGI-3 launch announcement: evaluates agentic intelligence via interactive reasoning environments. 100% solvable by humans with no training, but all frontier AI reasoning models score under 1%.