Malign initializations are more robust when the model can think better in the reasoning language than in the output language
One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here’s a fairly typical result from our previous research:We train a (reasoning) malign init to sandbag on some inputs.We SFT the model on responses to simple questions, generated by a differe...
Read full article →