Confirming Claims of Superposition and Adversarial Examples in Toy Models
This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship.tl;dr:We reproduce all three core claims of Stevinson et al. from their toy classifier setting:PGD attacks against toy models generally agree with theoretically optimal solutions.Toy models without superposition are less vulnerable to attacks. Robustness falls monotonically as superposition increases.Attacks transfer between independently-tr...
Read full article →