we are now benchmarking our models on novel frontier research, via https://t.co/2XmndVes5F. of 10 m...
By @gdb
Building on yesterday's Reddit discussion about the First Proof results, Greg Brockman announces that OpenAI is benchmarking models on novel frontier research via 'First Proof' - their model found likely correct solutions to at least 6 of 10 unpublished math research problems in a week.