Overtly misaligned trajectories score highly in RL.

·LessWrong··

It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas KwaConstellation vs MIRI vs Reality Current AI agents often behave overtly egregiously misaligned. By “overt”, I mean that a human, reading the transcript, would say “The agent is obviously acting in direct opposition to the specification, user intent, and any common-sense understanding of good behaviour.” The reason, it se...

Read full article →

Related Articles

Two-tier encryption in the UK
ReturnoftheHack · Hacker News · 14h ago
F-Droid 2.0
daveoc64 · Hacker News · 9h ago
Creatine uptake enhances antitumor immunity
lormayna · Hacker News · 6h ago
Italian parliament votes for return to nuclear energy
geox · Hacker News · 1d ago
Google’s Project Suncatcher to put ML infrastructure in space
xnx · Hacker News · 11h ago