Steering Role Confusion

·LessWrong··

Summary: The recent work A Mechanistic Explanation of Prompt Injections puts forward a theory that prompt injection attacks succeed due to role confusion: models primarily infer which role produced a piece of text from forgeable cues rather than its role tags.The paper's experiments establish that such cues shift latent role representation, and also increase downstream compliance with prompt injections.However, they do not show that shifts in latent role representation directly influence complia...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 20h ago
Asahi Linux Now Officially Supports Apple M3 Macs – With Caveats
mdp2021 · Hacker News · 3h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 21h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago