MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses

·LessWrong··

A group of friends and I spent the last several months running an experiment in our free time to determine if a MUD would be a suitable environment for benchmarking and evaluating LLMs. The results of the experiment were not what we expected. The main surprise was that the model rankings were extremely sensitive to the individual components of each score, especially so for those which depended on an LLM classifier. The overall data was too broad to help us understand which model was most impacte...

Read full article →

Related Articles

Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 1d ago
The Art of 64-bit Assembly
0x54MUR41 · Hacker News · 12h ago
Tailscale didn't stop the Hugging Face intrusion
bluehatbrit · Hacker News · 1d ago
CISA Alert: Water Sector PLC Targeting
speckx · Hacker News · 8h ago
Postmortem for Kernel Soundness Bug #14576
juhopitk · Hacker News · 8h ago