Can we use steering vectors to suppress reward-hacking? Somewhat

·LessWrong··

Can steering vectors drive gradient routing? Yes, but not in realistic reward hacking environments, they are not precise enough classifiers of hacky vs clean solutions. Instead, can we use a steering vector to initialise adapters so that gradient routing happens without a classifier, and we get automatic seperation of hacky and clean gradients? Partly! This init approach suppressed 70% of hacking by absorbing gradients into the hacky initialised adapter. This is not as good as the prior approach...

Read full article →

Related Articles

England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 23h ago
London Underground begins scanning passengers' faces
BlueBerry2001 · Hacker News · 1d ago
Beef and dairy drive 41% of biodiversity damage linked to global farmland
robtherobber · Hacker News · 2h ago
Stealing Reasoning Traces from Proprietary LLM APIs
quantumgarbage · Hacker News · 22h ago
CFTC declares market emergency, orders Kalshi to continue to operate in New York
michaefe · Hacker News · 11h ago