llms exposed to a gcg trigger optimised for shannon entropy will randomly choose a persona and stay in it

·LessWrong··

Background/introI did this work under Suvajit Majumder's supervision as part of Eleuther AI's SOAR program. If you don't know what GCG is in the context of AI jailbreaking and Shannon entropy, I recommend asking your favourite AI before reading further. In our SOAR stream, we've been working on subliminal prompting. This is trying to non obviously change a model's preferences through prompting. For example, one of our goals is getting LLMs to produce insecure code without obviously telling it to...

Read full article →

Related Articles

Asahi Linux on M3
mdp2021 · Hacker News · 9h ago
The car industry A/B tested selling a car with and without CarPlay
gumby · Hacker News · 3h ago
Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 1d ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 2h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 1d ago