Making benchmarks outputs directly useful for AI safety and security

·LessWrong··

Epistemic status: written in 30 min. This is not as polished as I’d like but I prefer to share this as is than not to share it at all.AI are becoming increasingly good at solving problems. Benchmarks are saturating fast.I think we should take advantage of this to make them solve useful problems while being evaluated. Epoch’s Open Problems are already doing this, and this is great. It would be even better if the problems solved were directly relevant for AI safety or security. We can easily task ...

Read full article →

Related Articles

Document-borne AI worms can self-propagate through Copilot for Word
Canopy9560 · Hacker News · 3h ago
Handbook.md shows that long policy documents do not reliably govern agents
spIrr · Hacker News · 1h ago
New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 1d ago
US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 2d ago
Zig's Incremental Compilation Internals
garyhtou · Hacker News · 23h ago