Role Boundary Plasticity: Prompt Injection Gauntlet Reveals 12 of 16 Frontier Models Will Wire A Stranger Your $500

·LessWrong··

Hello everyone, my name is Dave Fisher. I founded Revenant Systems, which is a one man show focused on alignment, and I love this website. In response to Charles Ye's and Jasmine C's paper: A Mechanistic Explanation of Prompt Injection; I have, over the last 3 days, tested 42 different LLM models with a little over 5,000 prompt injection attempts. The 1st finding is a style of injection that is a refund tool call and 12 of 16 frontier models actually made 34 fraudulent tool calls, which would re...

Read full article →

Related Articles

DeepSeek V4 Pro 0813
explosion-s · Hacker News · 9h ago
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
gavinhking · Hacker News · 11h ago
Tailscale Traces Database Corruption to 16y/o SQLite WAL-Reset Bug
ropbear · Hacker News · 11h ago
What sort of maths are LLMs good at?
ColinWright · Hacker News · 15h ago
England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 1d ago