MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses

·LessWrong··

A group of friends and I spent the last several months running an experiment in our free time to determine if a MUD would be a suitable environment for benchmarking and evaluating LLMs. The results of the experiment were not what we expected. The main surprise was that the model rankings were extremely sensitive to the individual components of each score, especially so for those which depended on an LLM classifier. The overall data was too broad to help us understand which model was most impacte...

Read full article →

Related Articles

We got admin access to Baseten's production GitHub in 25 minutes
bearsyankees · Hacker News · 9h ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 8h ago
Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 15h ago
America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 12h ago
How much oil-market buffer is left?
mcone · Hacker News · 8h ago