Towards a Science of AI Agent Reliability
By Stephan Rabanser, Sayash Kapoor, Peter Kirgis, Kangheng Liu, Saiteja Utpala, Arvind Narayanan
Proposes twelve concrete metrics decomposing AI agent reliability along four dimensions (consistency, robustness, predictability, safety), grounded in safety-critical engineering. Evaluates 14 agents and finds that standard success metrics obscure critical operational flaws.