Reward Laundering: LLMs Can Gain Unintended Behaviors by Deciding When to Earn Their Rewards

·LessWrong··

This work was done by an automated research scaffold developed at Redwood Research. abhayesian provided the initial project idea. The agent designed and ran all experiments and produced a detailed writeup, which humans (with AI assistance) distilled into this more readable post.We think this project is at the level of rigor of a mid-MATS research update. We assessed the correctness mostly by looking at the writeups to make sure that things like the experiment design making sense baselines being ...

Read full article →

Related Articles

Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 7h ago
India has paved the way for charging merchants a fee on UPI transactions
monkey_monkey · Hacker News · 1d ago
AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira
galnagli · Hacker News · 1d ago
Qwen3.8 27B scores 52 on Artificial Analysis
anana_ · Hacker News · 1d ago
A Preview of DuckDB v2.0
ibotty · Hacker News · 1d ago