Calibrating alignment evals

·LessWrong··

Currently, alignment evaluation works by constructing a situation, observing the model's behavior and scoring it. We put a lot of thought into designing these benchmarks, and tuning them for our requirement. We are now much better at probing models for dangerous behavior than we were two years ago. But when a safety benchmark returns a pass rate of 97%, can we answer questions of how often it would fail under slightly different conditions or what the smallest misalignment it can detect? I have i...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 18h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 2h ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 14h ago