Training On Interpretability Probes Is Bad In Proportion To How Contingent The Features They Rely On Are

·LessWrong··

People spend a lot of words playing tug of war over whether or not it's reasonable to train against interpretability methods. The anti case goes something like "training based on interpreted features trains against Interpretability itself more than it trains against whatever features you're detecting". There are cases where we should expect this to be true and cases where we should expect this to be not true. It basically comes down to how much the model can encrypt/obfuscate the relevant featur...

Read full article →

Related Articles

Omarchy: Any User Process Can Escalate to Root
trap0xcc · Hacker News · 4h ago
Bug Blindness
davidmckenna · Hacker News · 19h ago
Hy4 preview
shenli3514 · Hacker News · 1d ago
Haiku R1/beta6 has been released
metrofun · Hacker News · 4h ago
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
catbird · Hacker News · 6h ago