Confidence Estimation in Automatic Short Answer Grading with LLMs

·ArXiv cs.CL··

arXiv:2605.00200v1 Announce Type: new Abstract: Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fine-tuning, while also enabling the generation of synthetic feedback for educational assessment. Despite these advances, LLM-based grading remains imperfect, making reliable confidence estimates essential for safe and effective human-AI collaboration in educational decision-making. In this work, we i...

Read full article →

Related Articles

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
donsupreme · Hacker News · 5mo ago
Accelerating Gemma 4: faster inference with multi-token prediction drafters
amrrs · Hacker News · 5mo ago
Harvard particle physicist Matthew Schwartz drops 36 papers authored with Claude
xqcgrek2 · Hacker News · 1d ago
An AI agent emailed researchers for help. It told us why
sbulaev · Hacker News · 12h ago
A couple million lines of Haskell: Production engineering at Mercury
unignorant · Hacker News · 5mo ago