Introducing GeneBench-Pro — testing whether models can handle the kind of judgment-heavy analysis th...
By @gdb
Greg Brockman introduces GeneBench-Pro, a benchmark testing judgment-heavy computational biology tasks that take human experts 20-40 hours, and highlights GPT-5.6 Sol as a big step forward.