We're talking past our models; or, How a model defined its "evil" vector as dread

·LessWrong··

SummaryWe train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector.To learn how the model interprets this steering vector, we then ask the model to a) respond in the style of this neologism, and b) explain it.Responses generated with the neologism are substantially more similar to the steering vector (larger projection values) than responses generated with the steering vector itself, while bei...

Read full article →

Related Articles

OpenAI's GPT-6 Astra on ARC-AGI-3
vignesh_warar · Hacker News · 8h ago
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
screm · Hacker News · 7h ago
GLP-1s are being linked to fewer serious infections, including TB
gumby · Hacker News · 5h ago
Pre-Release of Polars 2.0
komape · Hacker News · 21h ago
Artificial beaver dams saw juvenile coho salmon survival rates go from 8% to 60%
speckx · Hacker News · 12h ago