We're talking past our models; or, How a model defined its "evil" vector as dread

·LessWrong··

SummaryWe train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector.To learn how the model interprets this steering vector, we then ask the model to a) respond in the style of this neologism, and b) explain it.Responses generated with the neologism are substantially more similar to the steering vector (larger projection values) than responses generated with the steering vector itself, while bei...

Read full article →

Related Articles

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
cl42 · Hacker News · 10h ago
Hacker wipes Romania's land registry database
speckx · Hacker News · 11h ago
Claude Fable produced a counterexample to the Jacobian Conjecture
loubbrad · Hacker News · 22h ago
Claude Code uses Bun written in Rust now
tosh · Hacker News · 1d ago
How we measured AI writing across arXiv, and where the measurement breaks
dopamine_daddy · Hacker News · 8h ago