Why Do Naive SFT Filters For Safety Properties Fail?

·LessWrong··

This is the fourth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The third post can be found here.Since SFT is the cause for many safety relevant properties, a natural strategy is to filter out rollouts from SFT that have undesirable properties. However, as we show in this section (and in forthcoming MATS work), SFT data filtering frequently works surprisingly poorly. In this post, we investigate hy...

Read full article →

Related Articles

google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 1d ago
Will There Be a 7G?
Betelbuddy · Hacker News · 12h ago
Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
theanonymousone · Hacker News · 8h ago
Navier-Stokes Announcement
rvz · Hacker News · 1d ago
Linux Zoom client proactively reading everything written to X11 clipboard
encyclopedism · Hacker News · 10h ago