A Calibration Benchmark for LLM Beliefs Across a Taxonomic Hierarchy by DanRKAlex

·Nuno Sempere··

TL;DR: I’ve built a few small cal­ibra­tion bench­marks to test whether LLMs rep­re­sent un­cer­tainty well, not just whether they’re ac­cu­rate. The most re­cent one found in­con­sis­ten­cies in some LLMs’ pre­dic­tions/​guesses in ad­di­tion to and ir­re­spec­tive of their de­vi­a­tion from the ground truth. The test is some­what un­der­pow­ered (n=54×6) but I still found the re­sults in­ter­est­ing.Back­ground: I’m an economist by train­ing (MS Iowa State) who moved into data sci­ence (MS UMi...

Read full article →

Related Articles

The Copier Line: What ForecastBench’s Market Scores Actually Measure. by Dominus
Dominus · Nuno Sempere · 5h ago
What Effective AI Capability Investment Looks Like in Low Resource Research Settings by Alejandra Carriero
Alejandra Carriero · Nuno Sempere · 23h ago
Thoughts on Taking OpenAI Foundation Funding by Jeff Kaufman 🔸
Jeff Kaufman 🔸 · Nuno Sempere · 1d ago
What do AI assistants tell users about tobacco harm reduction? by Kristof Redei
Kristof Redei · Nuno Sempere · 1d ago
Nearly 8 Billion Male Chicks Culled Per Year: A Transparent Global Estimate (Faunalytics) by JLRiedi
JLRiedi · Nuno Sempere · 1d ago