Terminal-Bench Leaderboard Rankings: Luck or Skill?
Terminal-Bench 2.0, one of the best agent benchmarks, runs agents on 89 tasks with at least 5 tries per task. The aggregate rankings are then published along with a confidence interval. What I wanted to know, however, was whether or not the point gap between any two given agents represented a statistically significant difference (or if it was just noise).My first finding was that if you look at pairs with adjacent rankings then 24/25 of them differ by such a small amount that it's essentially st...
Read full article →