Research note on negated reward hacking

·LessWrong··

This work was done as part of the BlueDot's Technical AI Safety Project Sprint and should be treated as an informal report of preliminary results done over a couple of days.The code is available on GitHub, the negated dataset and model checkpoints are available on HuggingFace, also the hack rollouts are available at this viewer.IntroductionEmergent misalignment (EM) is a phenomenon where narrow fine-tuning can cause broadly misaligned behaviors in current LLMs. Both Anthropic (MacDiarmid et al.,...

Read full article →

Related Articles

Italian parliament votes for return to nuclear energy
geox · Hacker News · 15h ago
Claude Opus 5.5
km144 · Hacker News · 1d ago
Linux support is coming to Snapdragon X2 Series
aaronday · Hacker News · 9h ago
Claude Code reads AGENTS.md only when telemetry is on [fixed]
pszypowicz · Hacker News · 20h ago
GPT-6 Astra has gained the ability to drive a car
plurby · Hacker News · 17h ago