Testing LLMs on Undergraduate Music Theory

·LessWrong··

I spent the past week designing a test that I hoped would serve as a benchmark. But LLMs are improving faster than I expected, and my devilishly hard questions turned out to be a cakewalk.Here are the results of testing five modern LLMs on undergraduate music theory:As can be seen, the LLMs performed remarkably well. Each scored a passing grade and GPT 5.6 Sol (the only premium model tested) scored a perfect 100%.With results like these, there’s no point in using this test as a benchmark going f...

Read full article →

Related Articles

Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 21h ago
I'm being cyberattacked by Tesla, Inc
robinpie · Hacker News · 5h ago
JetKVM Mini
taubek · Hacker News · 15h ago
google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 1d ago
Revolut confirms customer data breach through fake government requests
tdrz · Hacker News · 13h ago