Fixed-weight models are adversarially vulnerable: hence misaligned

·LessWrong··

This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure.Boundaries in concept spaceTo serve any purpose whatsoever, an AI will have to draw boundaries inside its world-model - to distinguish world A from world B, and reach some comparison between them.If we want the AI to follow our goals and values, we want it to be able t...

Read full article →

Related Articles

Nissan's third generation e-POWER powertrain
mroche · Hacker News · 18h ago
Nvidia wants to put a watchdog chip next to every AI agent
jonbaer · Hacker News · 5h ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 3d ago
MicroLLM Lab – Try 7 tiny LLM's in the browser
logicallee · Hacker News · 2h ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 3d ago