Watching the Wrong Thing: an interpretability audit of two audio deepfake detectors by Gavin_Lee
Project report. Full code, artifacts, and the append-only experiment log are in the project repository. Can you tell, by looking inside a model, when it is being fooled? I started this project expecting the answer to be yes. The plan was to build an internal-representation monitor for adversarial audio and show that internals were more trustworthy than the model’s own confidence. That headline did not survive contact with the experiments. The project en...
Read full article →