Are We Guarding Against Backdoors Or Failing To Notice Them? (Part 1 / 6)

·LessWrong··

This post serves to argue that backdooring evaluations are prone to failures stemming from triggers never reaching models.In backdooring literature, there is a common workflow. Outputs are evaluated on inputs that contain triggers. The outcome is thus clear. If the model did not display the backdoor despite ingesting the trigger, it is considered robust[1].This process, as described, skips a critical step; AI safety researchers and eval builders may benefit from carefully evaluating if triggers ...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 1d ago
The Mathocalypse
6bitquant · Hacker News · 13h ago
Shipping JPEG XL in Chrome
AshleysBrain · Hacker News · 21h ago
Navier–Stokes Lost in Translation
nill0 · Hacker News · 17h ago
Port of the TypeScript compiler, checker and lsp to Rust, by LLM
jcbhmr · Hacker News · 8h ago