Defeating Introspection Adapters (and Why Threat Models Matter)

·LessWrong··

We demonstrated an attack against Introspection Adapters (Shenoy et al., 2026), a technique for detecting malicious fine-tunes. Long story short: an attacker who controls model weights can apply a cheap, output-preserving transform that relocates the basis the auditor was calibrated against. This defeats the auditor with no observable change in model behavior. 📄 Paper💻 CodeAfter we discovered this attack, asked Keshav Shenoy (the first author) what he thought. It turned out that their team had...

Read full article →

Related Articles

New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 7h ago
A walk through of the DeltaNet family of linear attention variants
AnhTho_FR · Hacker News · 4h ago
GrapheneOS Defends Data-Wiping Function That Blocked US Border Search
pseudolus · Hacker News · 5h ago
US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 1d ago
Zig's Incremental Compilation Internals
garyhtou · Hacker News · 4h ago