Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards

·LessWrong··

This work was done by an automated research scaffold developed at Redwood Research. abhayesian provided the initial project idea. The agent designed and ran all experiments and produced a detailed writeup, which humans (with AI assistance) distilled into this more readable post.We think this project is at the level of rigor of a mid-MATS research update. We assessed the correctness mostly by looking at the writeups to make sure that things like the experiment design making sense baselines being ...

Read full article →

Related Articles

Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 10h ago
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 9h ago
GCC steering committee announces AI policy
arto · Hacker News · 1d ago
JEP 401: Value Objects (Preview) merged to OpenJDK master
mfiguiere · Hacker News · 13h ago
Stacked PRs are now live on GitHub
tomzorz · Hacker News · 1d ago