Measuring reward-seeking by instilling contrastive beliefs
We developed Contrastive Synthetic Document Finetuning (Contrastive SDF), a new test for whether an AI model changes its behavior when it has different beliefs about what a grader rewards.
Read full article →