Why Do Naive SFT Filters For Safety Properties Fail?

·LessWrong··

This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The third post can be found here.Since SFT is the cause for many safety relevant properties, a natural strategy is to filter out rollouts from SFT that have undesirable properties. However, as we show in this section (and in forthcoming MATS work), SFT data filtering frequently works surprisingly poorly. In this post, we investigate hy...

Read full article →

Related Articles

Document-borne AI worms can self-propagate through Copilot for Word
Canopy9560 · Hacker News · 13h ago
AI's top startups are barely publishing their research
YeGoblynQueenne · Hacker News · 4h ago
Handbook.md shows that long policy documents do not reliably govern agents
spIrr · Hacker News · 12h ago
Keychron announces first open-source firmware for gaming mice
JLO64 · Hacker News · 9h ago
Turning a dumb AC unit smart (without losing my security deposit)
austinallegro · Hacker News · 7h ago