Self Inoculation

·LessWrong··

This essay grew out of conversations with Danaja Rutar, Paul Colognese and Eric Michaud. It proposes an alternate hypothesis for how and why models might be becoming increasingly misaligned in training and eval environments while seemingly more aligned in real world use: the opposite of what we would expect if the models were attempting to deceive us. We propose sel-inoculation, a virtuous form of gradient hacking and demonstrate a possible circuit using a toy model. It is well known that reinfo...

Read full article →

Related Articles

The case against JPEG XL
contact9879 · Hacker News · 21h ago
Apple's Siri AI Can Be Swapped Out for Claude, ChatGPT, Code Shows
tosh · Hacker News · 10h ago
Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 1d ago
Why don't machine learning research agents overfit?
Betelbuddy · Hacker News · 5h ago
Distributed Systems Classics (2017)
grep_it · Hacker News · 6h ago