Three years of progress in 500 lines of code

·LessWrong··

TL;DRThere is some consensus that LLMs are bad at hard-to-verify tasks. The question is whether models are getting better at them over time. As a motivating example, I gave one research-reproduction task to 12 models spanning three years of progress, to illustrate how (1) what looked like an emergent capability was a predictable trend, visible years earlier if progress was measured with granularity, and (2) for the earliest models, building a verifiable check would have been close to impossible:...

Read full article →

Related Articles

US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 11h ago
Oracle bans AI-generated code from OpenJDK
delduca · Hacker News · 4h ago
AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 1d ago
Adults over 65 will outnumber children by 2029
brandonb · Hacker News · 6h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 1d ago