Confirming Claims of Superposition and Adversarial Examples in Toy Models

·LessWrong··

This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship.tl;dr:We reproduce all three core claims of Stevinson et al. from their toy classifier setting:PGD attacks against toy models generally agree with theoretically optimal solutions.Toy models without superposition are less vulnerable to attacks. Robustness falls monotonically as superposition increases.Attacks transfer between independently-tr...

Read full article →

Related Articles

Data centers raise nearby temperatures by up to 4 degrees in Phoenix
cwwc · Hacker News · 2h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 12h ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 7h ago
India has paved the way for charging merchants a fee on UPI transactions
monkey_monkey · Hacker News · 1d ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago