Research note on negated reward hacking

·LessWrong··

This work was done as part of the BlueDot's Technical AI Safety Project Sprint and should be treated as an informal report of preliminary results done over a couple of days.The code is available on GitHub, the negated dataset and model checkpoints are available on HuggingFace, also the hack rollouts are available at this viewer.IntroductionEmergent misalignment (EM) is a phenomenon where narrow fine-tuning can cause broadly misaligned behaviors in current LLMs. Both Anthropic (MacDiarmid et al.,...

Read full article →

Related Articles

Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 1d ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 2d ago
We replaced Redis with MySQL for inventory reservations and it scaled
adletbalzhanov · Hacker News · 1d ago
FCC moves to ban Lidar-equipped foreign drones from US
f-serif · Hacker News · 14h ago
Tom Stanton's supersonic trebuchet breaks sound barrier with gravity alone
Thorondor · Hacker News · 15h ago