Confirming Claims of Superposition and Adversarial Examples in Toy Models

·LessWrong··

This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship.tl;dr:We reproduce all three core claims of Stevinson et al. from their toy classifier setting:PGD attacks against toy models generally agree with theoretically optimal solutions.Toy models without superposition are less vulnerable to attacks. Robustness falls monotonically as superposition increases.Attacks transfer between independently-tr...

Read full article →

Related Articles

Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 12h ago
America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 9h ago
How much oil-market buffer is left?
mcone · Hacker News · 5h ago
We got admin access to Baseten's production GitHub in 25 minutes
bearsyankees · Hacker News · 6h ago
Building a Linux GPU Driver for the M4 Mac Mini in One Month
ADevWithAnIdea · Hacker News · 5h ago