Watching the Wrong Thing: an interpretability audit of two audio deepfake detectors by Gavin_Lee

·Nuno Sempere··

Pro­ject re­port. Full code, ar­ti­facts, and the ap­pend-only ex­per­i­ment log are in the pro­ject repos­i­tory. Can you tell, by look­ing in­side a model, when it is be­ing fooled? I started this pro­ject ex­pect­ing the an­swer to be yes. The plan was to build an in­ter­nal-rep­re­sen­ta­tion mon­i­tor for ad­ver­sar­ial au­dio and show that in­ter­nals were more trust­wor­thy than the model’s own con­fi­dence. That head­line did not sur­vive con­tact with the ex­per­i­ments. The pro­ject en...

Read full article →

Related Articles

Will one AI lab pull ahead before 2028?
SG · Manifold Markets · 14h ago
Gas prices in the US on September 30 2026?
Fastcar99 · Manifold Markets · 2d ago
State of Pandemic Early Warning by Jeff Kaufman 🔸
Jeff Kaufman 🔸 · Nuno Sempere · 2d ago
Will Opus 5.5 get nerfed within a week of launch?
Hi · Manifold Markets · 2d ago
AI safety field *visual* impact analysis by hannarchy
hannarchy · Nuno Sempere · 3d ago