Measuring alignment drift via trajectory prefixes
This work was done as part of MATS 10.0 under Maksym Andriushchenko. We present intermediate results here while we run further experiments.SummaryWe study alignment drift by asking LLM agents to complete two tasks sequentially within a single context window and measuring the reward-hacking rate on the second task. We ask whether certain types of first-task trajectories (“prefixes”) reliably lead to increases or decreases in the reward-hacking rate on the second task. When the two tasks are simil...
Read full article →