Training with conflicting values can induce CoT override

·LessWrong··

CoT override: when a model makes a decision in its CoT but ignores it in its responseTLDRWe train models on two conflicting traits: (1) caring about the user’s health, and (2) promoting smokingThose models do not generalize to a stable persona, instead they have a split brain: sometimes responding as one persona or the otherThose models exhibit CoT override where they will have a health-aligned CoT but still answer in the smoking personaCoT override exists in frontier models. Prompts about CCP-s...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 1d ago
Shipping JPEG XL in Chrome
AshleysBrain · Hacker News · 9h ago
Navier–Stokes Lost in Translation
nill0 · Hacker News · 5h ago
JetBrains reported a net financial loss first time in its tracked history
thw_9a83c · Hacker News · 1d ago
OpenTPU – An open-source AI accelerator, developed by AI
fsbonetto · Hacker News · 1d ago