Fixed-weight models are adversarially vulnerable: hence misaligned
This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure.Boundaries in concept spaceTo serve any purpose whatsoever, an AI will have to draw boundaries inside its world-model - to distinguish world A from world B, and reach some comparison between them.If we want the AI to follow our goals and values, we want it to be able t...
Read full article →