2B scoring model flags out-of-domain misalignment, suggesting specialist judges have potential for audits

·LessWrong··

TL;DRSome evidence that narrow ‘specialist’ models could be useful as part of deployed model misalignment audits, complementing larger frontier auditing agents and offering potential cost, discrimination and transparency benefits.A Gemma 2B specialist judge trained on Betley et al 2025b (‘Betley’) code examples is able to distinguish between responses provided by insecure-fine-tuned ‘misaligned’ and ‘secure-fine-tuned’ models responding to out-of-domain general safety prompts (‘ICEBERG’ - detail...

Read full article →

Related Articles

Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 6h ago
Moderna reports first positive Phase 3 for mRNA neoantigen therapy in melanoma
heydenberk · Hacker News · 4h ago
Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 1d ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 18h ago
The Mojo language (by Modular, now Qualcomm) is now open-source
flaburgan · Hacker News · 10h ago