Do VPD's Explanations Aggregate? An Audit of the Released Decomposition
TL;DR: adVersarial Parameter Decomposition (VPD) decomposes model weights into simple components and then labels each component as "needed here" or "safe to remove here" for each token of the input. These labels make up the "explanation" of that input. A core aspiration of VPD is that inputs' explanations can aggregate without changing the model's outputs. On this front, the authors themselves write, "It remains unclear whether our current decomposition is sufficiently adversarially robust for t...
Read full article →