Malign initializations are more robust when the model can think better in the reasoning language than in the output language

·LessWrong··

One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here’s a fairly typical result from our previous research:We train a (reasoning) malign init to sandbag on some inputs.We SFT the model on responses to simple questions, generated by a differe...

Read full article →

Related Articles

We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 6h ago
Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 6h ago
Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 14h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 5h ago
Decompiling a Nintendo 64 game in 84 days
knackers · Hacker News · 9h ago