A Theoretical Game of Attacks via Compositional Skills

·ArXiv cs.CL··

arXiv:2605.01034v1 Announce Type: new Abstract: As large language models grow increasingly capable, concerns about their safe deployment have intensified. While numerous alignment strategies aim to restrict harmful behavior, these defenses can still be circumvented through carefully designed adversarial prompts. In this work, we introduce a theoretical framework that formalizes a game between an attacker and a defender. Within this framework, we design a theoretical best-response attack strategy...

Read full article →

Related Articles

Reducing synthetic markers makes some SDF false facts linearly indistinguishable from pretraining-acquired knowledge
Jason Zeng · LessWrong · 39m ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 15d ago
Continual learning might make your blocking monitors nearly useless
Alex Mallen · Alignment Forum · 9d ago
Latent reasoning architectures would likely undermine CoT, our strongest oversight tool
Lukas Finnveden · Redwood Research · 10d ago