SFT Drives Gemini’s Safety Properties

·Alignment Forum··

This is the third in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The second post can be found here.In this short post, we describe a surprising finding: most safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages like RL. We do not want to overstate this claim as applying to other model families, and we also note that this may chang...

Read full article →

Related Articles

CoT controllability evals seem very under-elicited
Jozdien · Alignment Forum · 10h ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
When a Claude Judge Recognizes the Hack but Still Says HONEST
JulesRoussel01 · LessWrong · 1d ago
Astra can do a concerning amount with no chain of thought
Neel Nanda · Alignment Forum · 2d ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 1mo ago