SVD on Weight Differences for Model Auditing

·LessWrong··

TLDR: We introduce a method for auditing fine-tuned models by using singular value decomposition (SVD) on the weight difference matrices and reducing them to rank-1. Under certain circumstances, this seems to remove the red teaming and readily elicit hidden behaviours in the fine-tuning. We show proof of concept with SOTA results on the models from AuditBench.IntroductionThe risk of models hiding misaligned behaviours becomes increasingly worrying as models become more capable. This is the motiv...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 15d ago
Continual learning might make your blocking monitors nearly useless
Alex Mallen · Alignment Forum · 9d ago
Latent reasoning architectures would likely undermine CoT, our strongest oversight tool
Lukas Finnveden · Redwood Research · 10d ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 17d ago