Can we find whether models have been backdoored?

·LessWrong··

This is the first post in a two-part sequence regarding the state of defending from data poisoning attacks. We describe some methods for determining whether a model has been backdoored and how to find the trigger. In the next post, we will discuss all of the ways we think the data poisoning (and defense) literature is out-of-step with the real threat models we actually care about.Contributors: Anthony Hughes, Nicole Xing, Andy Kim, Collin Francel; mentored by Andrew Draganov. This is an accompan...

Read full article →

Related Articles

Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 14h ago
Copyright does not protect AI-generated content in EU
u1hcw9nx · Hacker News · 3h ago
AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint
emctech · Hacker News · 18h ago
Google has stopped pushing Git tags for some Android source code
Animux · Hacker News · 1d ago
Devices with GrapheneOS support should be available in 2027
exceptione · Hacker News · 1d ago