Three years of progress in 500 lines of code

·LessWrong··

TL;DRThere is some consensus that LLMs are bad at hard-to-verify tasks. The question is whether models are getting better at them over time. As a motivating example, I gave one research-reproduction task to 12 models spanning three years of progress, to illustrate how (1) what looked like an emergent capability was a predictable trend, visible years earlier if progress was measured with granularity, and (2) for the earliest models, building a verifiable check would have been close to impossible:...

Read full article →

Related Articles

NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 9h ago
What happened to the Snowden archive
EXHades · Hacker News · 1d ago
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 1d ago
Ask HN: Is it impossible to disable Siri on macOS 27?
semidror · Hacker News · 15h ago
HERMES radio enables voice and data communication over vast distances
SamuraiLion · Hacker News · 12h ago