Can we find whether models have been backdoored?

·LessWrong··

This is the first post in a two-part sequence regarding the state of defending from data poisoning attacks. We describe some methods for determining whether a model has been backdoored and how to find the trigger. In the next post, we will discuss all of the ways we think the data poisoning (and defense) literature is out-of-step with the real threat models we actually care about.Contributors: Anthony Hughes, Nicole Xing, Andy Kim, Collin Francel; mentored by Andrew Draganov. This is an accompan...

Read full article →

Related Articles

Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 12h ago
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 19h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 16h ago
A 40ms Go garbage collector pause caused by swap
shellpipe · Hacker News · 6h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 1d ago