2B scoring model flags out-of-domain misalignment, suggesting specialist judges have potential for audits

burnssa·LessWrong·Community·May 14, 2026

TL;DRSome evidence that narrow ‘specialist’ models could be useful as part of deployed model misalignment audits, complementing larger frontier auditing agents and offering potential cost, discrimination and transparency benefits.A Gemma 2B specialist judge trained on Betley et al 2025b (‘Betley’) code examples is able to distinguish between responses provided by insecure-fine-tuned ‘misaligned’ and ‘secure-fine-tuned’ models responding to out-of-domain general safety prompts (‘ICEBERG’ - detail...

Read full article →

2B scoring model flags out-of-domain misalignment, suggesting specialist judges have potential for audits

Related Articles