TASTE: Can AI Models Judge AI Safety Research Proposals?

·LessWrong··

tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers. Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement): a discussion stage in which researchers talk through disagreements before revising their scores, and filtering researchers’ labels for self-reported "strong" confide...

Read full article →

Related Articles

GLM-5.3 is now open-weight
jeudesprits · Hacker News · 5h ago
EPA says power for data centers can sidestep pollution laws
Levitating · Hacker News · 7h ago
Pentagon's blacklisting of Anthropic was unlawful, US judge rules
softwaredoug · Hacker News · 9h ago
Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 1d ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 1d ago