Malign initializations are more robust when the model can think better in the reasoning language than in the output language

·LessWrong··

One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here’s a fairly typical result from our previous research:We train a (reasoning) malign init to sandbag on some inputs.We SFT the model on responses to simple questions, generated by a differe...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 3h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 4h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago
Why are European countries moving their gold out of North America?
ranit · Hacker News · 18h ago