Do VPD's Explanations Aggregate? An Audit of the Released Decomposition

·LessWrong··

TL;DR: adVersarial Parameter Decomposition (VPD) decomposes model weights into simple components and then labels each component as "needed here" or "safe to remove here" for each token of the input. These labels make up the "explanation" of that input. A core aspiration of VPD is that inputs' explanations can aggregate without changing the model's outputs. On this front, the authors themselves write, "It remains unclear whether our current decomposition is sufficiently adversarially robust for t...

Read full article →

Related Articles

Livenerf: Has Opus 5.5 been nerfed yet?
bryan0 · Hacker News · 18h ago
Vermont replacing power plants with home batteries
devonnull · Hacker News · 22h ago
How Delhi cut electricity loss from 50 to 5 percent
rbanffy · Hacker News · 1d ago
US sanctions force The Netherlands off Microsoft and toward alternative NixOS
mywacaday · Hacker News · 1d ago
NASA asked several former SR-71A staffers to help secret restart
ilamont · Hacker News · 1d ago