Measuring reward-seeking by instilling contrastive beliefs

·Hacker News··

We developed Contrastive Synthetic Document Finetuning (Contrastive SDF), a new test for whether an AI model changes its behavior when it has different beliefs about what a grader rewards.

Read full article →

Related Articles

Apple defeats liability for not scanning iCloud for CSAM
speckx · Hacker News · 9h ago
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
piotrgrabowski · Hacker News · 1h ago
New US homeownership measure puts people first
throw0101a · Hacker News · 12h ago
FreeInk: Open ecosystem for e-readers
FriedPickles · Hacker News · 5h ago
Hacker wipes Romania's land registry database
speckx · Hacker News · 1d ago