Measuring alignment drift via trajectory prefixes

·LessWrong··

This work was done as part of MATS 10.0 under Maksym Andriushchenko. We present intermediate results here while we run further experiments.SummaryWe study alignment drift by asking LLM agents to complete two tasks sequentially within a single context window and measuring the reward-hacking rate on the second task. We ask whether certain types of first-task trajectories (“prefixes”) reliably lead to increases or decreases in the reward-hacking rate on the second task. When the two tasks are simil...

Read full article →

Related Articles

What is it like to be a neural net?
David Balduzzi · LessWrong · 23m ago
Agents let AI safety share experiments hourly, not just papers monthly
Jason Fantl · LessWrong · 27m ago
Constraining the capacity of physical side channels for AI verification and security
emlynsg · LessWrong · 35m ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 35m ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago