A Calibration Benchmark for LLM Beliefs Across a Taxonomic Hierarchy by DanRKAlex

·Nuno Sempere··

TL;DR: I’ve built a few small cal­ibra­tion bench­marks to test whether LLMs rep­re­sent un­cer­tainty well, not just whether they’re ac­cu­rate. The most re­cent one found in­con­sis­ten­cies in some LLMs’ pre­dic­tions/​guesses in ad­di­tion to and ir­re­spec­tive of their de­vi­a­tion from the ground truth. The test is some­what un­der­pow­ered (n=54×6) but I still found the re­sults in­ter­est­ing.Back­ground: I’m an economist by train­ing (MS Iowa State) who moved into data sci­ence (MS UMi...

Read full article →

Related Articles

How much should we worry about the pneumonic plague lableak in Siberia? by Drew Spartz
Drew Spartz · Nuno Sempere · 2h ago
Will the Irkutsk Anti-Plague institute incident turn into an outbreak?
Anthem · Manifold Markets · 1d ago
How many launches will SpaceX have in October?
Fastcar99 · Manifold Markets · 2d ago
What is recursive self-improvement, and what would it mean to ban it? by sarahhw
sarahhw · Nuno Sempere · 2d ago
State of the Field: AI for Epistemics and Coordination by Ben_N
Ben_N · Nuno Sempere · 2d ago