We're talking past our models; or, How a model defined its "evil" vector as dread

·LessWrong··

SummaryWe train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector.To learn how the model interprets this steering vector, we then ask the model to a) respond in the style of this neologism, and b) explain it.Responses generated with the neologism are substantially more similar to the steering vector (larger projection values) than responses generated with the steering vector itself, while bei...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 19h ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 3h ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
Meta Files Patent for Facial Recognition, Automatic Recording of People
DeepLogin · Hacker News · 14h ago