Are We Guarding Against Backdoors Or Failing To Notice Them? (Part 1 / 6)

·LessWrong··

This post serves to argue that backdooring evaluations are prone to failures stemming from triggers never reaching models.In backdooring literature, there is a common workflow. Outputs are evaluated on inputs that contain triggers. The outcome is thus clear. If the model did not display the backdoor despite ingesting the trigger, it is considered robust[1].This process, as described, skips a critical step; AI safety researchers and eval builders may benefit from carefully evaluating if triggers ...

Read full article →

Related Articles

Coconut oil jet fuel matches kerosene's efficiency in engine tests
mdp2021 · Hacker News · 16h ago
Slovakia finds Russian backdoor in traffic speed cameras
dredmorbius · Hacker News · 17h ago
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
ed-is-ai · Hacker News · 15h ago
There's no reason for software to be slow anymore
Jach · Hacker News · 2d ago
How Complex Systems Fail (1998)
shortcrct · Hacker News · 16h ago