Watching the Wrong Thing: an interpretability audit of two audio deepfake detectors by Gavin_Lee

·Nuno Sempere··

Pro­ject re­port. Full code, ar­ti­facts, and the ap­pend-only ex­per­i­ment log are in the pro­ject repos­i­tory. Can you tell, by look­ing in­side a model, when it is be­ing fooled? I started this pro­ject ex­pect­ing the an­swer to be yes. The plan was to build an in­ter­nal-rep­re­sen­ta­tion mon­i­tor for ad­ver­sar­ial au­dio and show that in­ter­nals were more trust­wor­thy than the model’s own con­fi­dence. That head­line did not sur­vive con­tact with the ex­per­i­ments. The pro­ject en...

Read full article →

Related Articles

Geoffrey Irving on how to solve alignment before superintelligence arrives by 80000_Hours
80000_Hours · Nuno Sempere · 1d ago
Will I see the total solar eclipse for at least ~5 seconds tomorrow?
Eli Goldfine · Manifold Markets · 2d ago
Will I solve an unsolved math problem with AI in August?
Bayesian · Manifold Markets · 2d ago
Coercion and Deception in AI-to-AI Management by Jonah Woodward
Jonah Woodward · Nuno Sempere · 2d ago
Aim at agency: institutional design under preference uncertainty, and why alignment needs it by act65
act65 · Nuno Sempere · 2d ago