Terminal-Bench Leaderboard Rankings: Luck or Skill?

·LessWrong··

Terminal-Bench 2.0, one of the best agent benchmarks, runs agents on 89 tasks with at least 5 tries per task. The aggregate rankings are then published along with a confidence interval. What I wanted to know, however, was whether or not the point gap between any two given agents represented a statistically significant difference (or if it was just noise).My first finding was that if you look at pairs with adjacent rankings then 24/25 of them differ by such a small amount that it's essentially st...

Read full article →

Related Articles

Data centers raise nearby temperatures by up to 4 degrees in Phoenix
cwwc · Hacker News · 2h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 12h ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 7h ago
India has paved the way for charging merchants a fee on UPI transactions
monkey_monkey · Hacker News · 1d ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago