UK's AI Security Institute finds standard benchmarks systematically underestimate what AI agents can actually do
By Matthias Bastian
The UK's AI Security Institute found that standard benchmarks systematically underestimate agent capabilities because they cap compute budgets. Increasing token budgets tenfold raised software-engineering success rates by about 25%, implying frontier progress is roughly 60% steeper than prior measurements showed.