Testing LLMs on Undergraduate Music Theory

·LessWrong··

I spent the past week designing a test that I hoped would serve as a benchmark. But LLMs are improving faster than I expected, and my devilishly hard questions turned out to be a cakewalk.Here are the results of testing five modern LLMs on undergraduate music theory:As can be seen, the LLMs performed remarkably well. Each scored a passing grade and GPT 5.6 Sol (the only premium model tested) scored a perfect 100%.With results like these, there’s no point in using this test as a benchmark going f...

Read full article →

Related Articles

Why is everyone trying to build a solid-state battery?
crescit_eundo · Hacker News · 6h ago
GCC steering committee announces AI policy
arto · Hacker News · 7h ago
AI's top startups are barely publishing their research
YeGoblynQueenne · Hacker News · 21h ago
Why Don't People Use Formal Methods? (2019)
Thom2503 · Hacker News · 6h ago
Document-borne AI worms can self-propagate through Copilot for Word
Canopy9560 · Hacker News · 1d ago