Terminal-Bench Leaderboard Rankings: Luck or Skill?

·LessWrong··

Terminal-Bench 2.0, one of the best agent benchmarks, runs agents on 89 tasks with at least 5 tries per task. The aggregate rankings are then published along with a confidence interval. What I wanted to know, however, was whether or not the point gap between any two given agents represented a statistically significant difference (or if it was just noise).My first finding was that if you look at pairs with adjacent rankings then 24/25 of them differ by such a small amount that it's essentially st...

Read full article →

Related Articles

Civilian plane crash in New Mexico tied to military GPS blocking
dzdt · Hacker News · 12h ago
Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 1d ago
Muse Code and Muse Spark 1.2
paulkrush · Hacker News · 4h ago
Ten advances in mathematics and theoretical computer science
milkshakes · Hacker News · 2d ago
Why Erdős Problems Are Falling to AI
pseudolus · Hacker News · 12h ago