What Are We Actually Benchmarking in Robot Manipulation?
By Tianchong Jiang, Xiangshan Tan, Samuel Wheeler, Luzhe Sun, Tewodros W. Ayalew, Matthew Walter
This paper critically examines robot manipulation benchmarks, identifying four failure modes (shortcut solvability, lack of statistical significance, creeping overfitting, data-source dependence) and proposing diagnostics for each. Auditing LIBERO, CALVIN, SimplerEnv, RoboCasa, and RoboTwin 2.0 reveals that popular benchmarks fail multiple diagnostics and a tiny probe can reach near-SOTA.