LaughBench
Introducing LaughBench. For a long time, I've considered the ability for AI models to tell novel, funny jokes that actually make people laugh to be a robust indicator of real general intelligence (as opposed to, say, coding tasks). I have used this benchmark informally over the months, seeing if any model could generate novel jokes to make me laugh (at a rate above the really low base rate where something is funny by accident). So far, the answer has always been no. As you can see from the chart...
Read full article →