Role Boundary Plasticity: Prompt Injection Gauntlet Reveals 12 of 16 Frontier Models Will Wire A Stranger Your $500

·LessWrong··

Hello everyone, my name is Dave Fisher. I founded Revenant Systems, which is a one man show focused on alignment, and I love this website. In response to Charles Ye's and Jasmine C's paper: A Mechanistic Explanation of Prompt Injection; I have, over the last 3 days, tested 42 different LLM models with a little over 5,000 prompt injection attempts. The 1st finding is a style of injection that is a refund tool call and 12 of 16 frontier models actually made 34 fraudulent tool calls, which would re...

Read full article →

Related Articles

Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 1d ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 1d ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 1d ago
Ask HN: Who's still keeping a DOS machine up because the business depends on it?
mlaux · Hacker News · 1d ago
DeepSeek Elastic Compute (DSec)
shenli3514 · Hacker News · 8h ago