Comment on Measuring Reward-Seeking by Instilling Contrastive Beliefs paper from mechanistic interpretability perspective

·LessWrong··

This is interesting research! https://alignment.openai.com/measuring-reward-seekingIt made me think of few overlapping hypotheses for what might be happening here, how did the grader behavior emerge at the pretraining and posttraining stages, which then gets shown in their evals and in production at test time:- Hypothesis 1: From pretraining the model already possesses tokens/representations/features/circuits of concepts such as tests, evaluators, success criteria, unit tests, oversight, etc., a...

Read full article →

Related Articles

LG to ban residential proxies from smart TV apps
DemiGuru · Hacker News · 18h ago
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
piotrgrabowski · Hacker News · 21h ago
Apple defeats liability for not scanning iCloud for CSAM
speckx · Hacker News · 1d ago
New US homeownership measure puts people first
throw0101a · Hacker News · 1d ago
FreeInk: Open ecosystem for e-readers
FriedPickles · Hacker News · 1d ago