In other words: The influence of prompt variation on alignment evals

·LessWrong··

If alignment evaluations are phrased differently, does model behaviour change? In this post, we describe our attempt to answer this question. This project was completed as a 5-day project sprint during the Capstone Week of ARENA 8.0.TL;DR: The data we collected were noisy! Almost every eval, model, and prompt variation direction yielded inconsistent results, making it difficult to draw clear conclusions. Nonetheless, our data suggest that prompt rephrasing can have a measurable impact on alignme...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 8h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 17h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 13h ago
GLM-5.3 Artificial Analysis Benchmarks
apitman · Hacker News · 3h ago