Cognitive Reasoning Diversity for Robust AI Juries

·LessWrong··

This project was done as part of BlueDot's Technical AI Safety Project Sprint under the mentorship of Jess Bergs.TL;DRResearchers have suggested that Human-AI juries may be more robust to judge hacking due to the complementarity of their orthogonal, uncorrelated blind spots In this exploratory project, these juries are simulated in silico with diverse cognitive reasoning strategies represented amongst judges to isolate, study, and validate the complementarity of their varied blind spots. With a ...

Read full article →

Related Articles

F-Droid 2.0
daveoc64 · Hacker News · 17h ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 19h ago
Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 22h ago
Toyota is taking the Corolla electric
cisc · Hacker News · 1d ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago