How useful is cross-domain generalization for training LLM monitors?

·LessWrong··

We study how well single-token classification training generalizes and find that:Control monitor performance can be improved by training on adjacent classification tasks (which is useful if we don't have high-quality in-domain data).We do not find strong benefits from classification-specialized models: training on instruction following after classification training not only keeps the uplift from classification training, but it also mitigates some generalization failures that arise from training ...

Read full article →

Related Articles

Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 11h ago
Pi 1.0
sergiotapia · Hacker News · 2d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 6h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago