Who else is steering?

·LessWrong··

This post raises a methodological question with regard to how to make sense of models' behaviors in response to activation steering anchored on a concept.Our discussion is limited to a particular kind of behaviors induced by activation steering, that of introspection studied by Lindsey (2026) (first accessed at here) and a range of investigations it inspires (see LW Post and citations in there). That said, the implication of our discussion applies to activation steering in general, mutatis mutan...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 1mo ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 1mo ago
An OpenAI model left notes about how to evade containment
Alex Mallen · Redwood Research · 1mo ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 1mo ago