In other words: The influence of prompt variation on alignment evals

·LessWrong··

If alignment evaluations are phrased differently, does model behaviour change? In this post, we describe our attempt to answer this question. This project was completed as a 5-day project sprint during the Capstone Week of ARENA 8.0.TL;DR: The data we collected were noisy! Almost every eval, model, and prompt variation direction yielded inconsistent results, making it difficult to draw clear conclusions. Nonetheless, our data suggest that prompt rephrasing can have a measurable impact on alignme...

Read full article →

Related Articles

LG to ban residential proxies from smart TV apps
DemiGuru · Hacker News · 23h ago
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
piotrgrabowski · Hacker News · 1d ago
Apple defeats liability for not scanning iCloud for CSAM
speckx · Hacker News · 1d ago
Everyone Should Know SIMD
WadeGrimridge · Hacker News · 7h ago
GigaToken: ~1000x faster Language model tokenization
syrusakbary · Hacker News · 7h ago