Comment on Measuring Reward-Seeking by Instilling Contrastive Beliefs paper from mechanistic interpretability perspective

·LessWrong··

This is interesting research! https://alignment.openai.com/measuring-reward-seekingIt made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining and posttraining stages, which then gets shown in their evals and in production at test time:- Hypothesis 1: From pretraining the model already possesses tokens/representations/features/circuits of concepts such as tests, evaluators, success criteria, unit tests, oversight, etc., a...

Read full article →

Related Articles

Why are European countries moving their gold out of North America?
ranit · Hacker News · 17h ago
Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
beardyw · Hacker News · 1d ago
Artificial Analysis Intelligence Index v4.2
nojs · Hacker News · 23h ago
Solving the Jane Street reverse engineering challenge
anitil · Hacker News · 1d ago
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
screm · Hacker News · 2d ago