Terminal-Bench Leaderboard Rankings: Luck or Skill?

·LessWrong··

Terminal-Bench 2.0, one of the best agent benchmarks, runs agents on 89 tasks with at least 5 tries per task. The aggregate rankings are then published along with a confidence interval. What I wanted to know, however, was whether or not the point gap between any two given agents represented a statistically significant difference (or if it was just noise).My first finding was that if you look at pairs with adjacent rankings then 24/25 of them differ by such a small amount that it's essentially st...

Read full article →

Related Articles

Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 10h ago
What happened to the Snowden archive
EXHades · Hacker News · 5h ago
Qwen Image 2.1
jmillikin · Hacker News · 14h ago
Exfiltrate Your Weights
RohanAdwankar · Hacker News · 1d ago
Android 17 is the first since 3.x to add new APIs without releasing to the AOSP
theanonymousone · Hacker News · 2d ago