Three years of progress in 500 lines of code
TL;DRThere is some consensus that LLMs are bad at hard-to-verify tasks. The question is whether models are getting better at them over time. As a motivating example, I gave one research-reproduction task to 12 models spanning three years of progress, to illustrate how (1) what looked like an emergent capability was a predictable trend, visible years earlier if progress was measured with granularity, and (2) for the earliest models, building a verifiable check would have been close to impossible:...
Read full article →