Measuring reward-seeking by instilling contrastive beliefs

·Hacker News··

We developed Contrastive Synthetic Document Finetuning (Contrastive SDF), a new test for whether an AI model changes its behavior when it has different beliefs about what a grader rewards.

Read full article →

Related Articles

Formalizing Fermat's Last Theorem
jlebar · Hacker News · 6h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 3h ago
Can AI design circuit boards yet?
iopapa · Hacker News · 5h ago
Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
beardyw · Hacker News · 18h ago
Solving the Jane Street reverse engineering challenge
anitil · Hacker News · 15h ago