Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

·Dan Luu··

We're going to look at three different kinds of benchmarks, one set of calculations for baseline numbers for performance "napkin math" estimates, one set of AI model evals, and one on car tires. To build my intuition for things, I like thinking about them before seeing the explanation, so these are presented with the benchmark information first and the explanation later in case you want to think about your answer before seeing my thoughts. 29. A friend of mine is reviewing performance orders of ...

Read full article →

Related Articles

Should I run plain Docker Compose in production in 2026?
pmig · Hacker News · 3mo ago
Bun is being ported from Zig to Rust
SergeAx · Hacker News · 3mo ago
Computer Use is 45x more expensive than structured APIs
palashawas · Hacker News · 3mo ago
Show HN: Tilde.run – Agent sandbox with a transactional, versioned filesystem
ozkatz · Hacker News · 3mo ago
RaTeX: KaTeX-compatible LaTeX rendering engine in pure Rust
atilimcetin · Hacker News · 3mo ago