Confirming Claims of Superposition and Adversarial Examples in Toy Models

·LessWrong··

This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship.tl;dr:We reproduce all three core claims of Stevinson et al. from their toy classifier setting:PGD attacks against toy models generally agree with theoretically optimal solutions.Toy models without superposition are less vulnerable to attacks. Robustness falls monotonically as superposition increases.Attacks transfer between independently-tr...

Read full article →

Related Articles

Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 1d ago
The Art of 64-bit Assembly
0x54MUR41 · Hacker News · 8h ago
Tailscale didn't stop the Hugging Face intrusion
bluehatbrit · Hacker News · 1d ago
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 1d ago
Postmortem for Kernel Soundness Bug #14576
juhopitk · Hacker News · 4h ago