A Mechanistic Explanation of Prompt Injection (and why you should study roles)

·LessWrong··

SummaryWe've been building a theory of how prompt injections work under the hood.We show it comes down to how LLMs perceive roles (the humble chat template tags).We use this theory to create new attacks, explain some weird mech interp results, and predict when attacks work.We also advocate for a new subfield focused on the science of roles, and sketch some unexplored new research problems.Work supported by CBAI and Cosmos. Another version of this post (with more inline colors) is here, and full ...

Read full article →

Related Articles

AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 10h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 12h ago
Nashville uses eminent domain to block data center near zoo
mapping365 · Hacker News · 1d ago
Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks
willcarkner · Hacker News · 15h ago
Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 2d ago