Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking

·LessWrong··

It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring[1], help us do better science on current models[2], and augment certain forms of alignment training[3]. Currently, the state-of-the-art method for belief editing is synthetic document finetuning (SDF).We test how well SDF works to inoculate a model against misalignment generalization from RL-induced reward hacking, by training models on documents framing reward...

Read full article →

Related Articles

America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 7h ago
How much oil-market buffer is left?
mcone · Hacker News · 3h ago
Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 10h ago
We got admin access to Baseten's production GitHub in 25 minutes
bearsyankees · Hacker News · 4h ago
Show HN: Capsule – Single-file web apps that save their data into SQLite
bashtian · Hacker News · 9h ago