Probing is not enough; a validity audit for any probe by Ratnaditya

·Nuno Sempere··

tl;drA probe can have ex­cel­lent AUROC and yet fail as a safety sig­nal. I au­dited 3 probes: a mon­i­tor­ing aware­ness probe leak­age ex­am­ple, a re­fusal di­rec­tion as pos­i­tive con­trol, and Apollo’s de­cep­tion probe as a pub­lished pro­to­col case study.In the leak­age case, the probe achieves AUROC 1.00, but falls to 0.50 when I re­move a sin­gle prompt tag. And re­mov­ing a ran­dom span of same length barely dents it. That re­sult uses only de­ter­minis­tic com­po­nents (avoids LLM j...

Read full article →

Related Articles

Ecological Collapse is not an Existential Risk by Tandena Wagner
Tandena Wagner · Nuno Sempere · 44m ago
The staggering welfare consequences of catch-up growth after AI by Rafael Proenca
Rafael Proenca · Nuno Sempere · 5h ago
Watching the Wrong Thing: an interpretability audit of two audio deepfake detectors by Gavin_Lee
Gavin_Lee · Nuno Sempere · 9h ago
Will Gemini 3.7 Flash be released by the end of August 2026?
JasonMendoza2008 · Manifold Markets · 1d ago
Geoffrey Irving on how to solve alignment before superintelligence arrives by 80000_Hours
80000_Hours · Nuno Sempere · 2d ago