llms exposed to a gcg trigger optimised for shannon entropy will randomly choose a persona and stay in it
Background/introI did this work under Suvajit Majumder's supervision as part of Eleuther AI's SOAR program. If you don't know what GCG is in the context of AI jailbreaking and Shannon entropy, I recommend asking your favourite AI before reading further. In our SOAR stream, we've been working on subliminal prompting. This is trying to non obviously change a model's preferences through prompting. For example, one of our goals is getting LLMs to produce insecure code without obviously telling it to...
Read full article →