Watching the Wrong Thing: an interpretability audit of two audio deepfake detectors by Gavin_Lee

·Nuno Sempere··

Pro­ject re­port. Full code, ar­ti­facts, and the ap­pend-only ex­per­i­ment log are in the pro­ject repos­i­tory. Can you tell, by look­ing in­side a model, when it is be­ing fooled? I started this pro­ject ex­pect­ing the an­swer to be yes. The plan was to build an in­ter­nal-rep­re­sen­ta­tion mon­i­tor for ad­ver­sar­ial au­dio and show that in­ter­nals were more trust­wor­thy than the model’s own con­fi­dence. That head­line did not sur­vive con­tact with the ex­per­i­ments. The pro­ject en...

Read full article →

Related Articles

Research Report: A genetic basis for potential nociception in cochineal bugs (Dactylopius coccus; Hemiptera: Dactylopiidae) by Meghan Barrett
Meghan Barrett · Nuno Sempere · 8h ago
We know who lives inside wild fish. We don’t know what happens in there by Pablo Sar
Pablo Sar · Nuno Sempere · 2d ago
Will Grok 4.7 be released on 9/11?
Eli Goldfine · Manifold Markets · 3d ago
Will I solve an unsolved math problem with AI in September?
Bayesian · Manifold Markets · 4d ago
Claude Fable 5.1+ (Anthropic) release date
Bayesian · Manifold Markets · 5d ago