Can we use steering vectors to suppress reward-hacking? Somewhat

·LessWrong··

Can steering vectors drive gradient routing? Yes, but not in realistic reward hacking environments, they are not precise enough classifiers of hacky vs clean solutions. Instead, can we use a steering vector to initialise adapters so that gradient routing happens without a classifier, and we get automatic seperation of hacky and clean gradients? Partly! This init approach suppressed 70% of hacking by absorbing gradients into the hacky initialised adapter. This is not as good as the prior approach...

Read full article →

Related Articles

Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 18h ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 1d ago
Ask HN: Who's still keeping a DOS machine up because the business depends on it?
mlaux · Hacker News · 19h ago
Excel now supports multiple values in a single cell
luispa · Hacker News · 18h ago
Toyota is taking the Corolla electric
cisc · Hacker News · 2d ago