Making benchmarks outputs directly useful for AI safety and security

·LessWrong··

Epistemic status: written in 30 min. This is not as polished as I’d like but I prefer to share this as is than not to share it at all.AI are becoming increasingly good at solving problems. Benchmarks are saturating fast.I think we should take advantage of this to make them solve useful problems while being evaluated. Epoch’s Open Problems are already doing this, and this is great. It would be even better if the problems solved were directly relevant for AI safety or security. We can easily task ...

Read full article →

Related Articles

google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 15h ago
Navier-Stokes Announcement
rvz · Hacker News · 15h ago
Measuring the sloppiness of code
doppp · Hacker News · 1d ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
Europe's "Less" Is Doing More Than Anyone Gives It Credit For
u1hcw9nx · Hacker News · 4h ago