Open-World Evaluations for Measuring Frontier AI Capabilities
By Sayash Kapoor, Peter Kirgis, Andrew Schwartz, Stephan Rabanser, J. J. Allaire, Rishi Bommasani, Harry Coppock, Magda Dubois, Gillian K Hadfield, Andrew B. Hall, Sara Hooker, Seth Lazar, Steve Newman, Dimitris Papailiopoulos, Shoshannah Tekofsky, Helen Toner, Cozmin Ududec, Arvind Narayanan
Advocates for 'open-world evaluations' - long-horizon, real-world tasks assessed through qualitative analysis rather than benchmark automation - as a complement to standard benchmarks. Introduces CRUX, a project for conducting such evaluations regularly, with initial findings on frontier models.