Synthetic document finetuning for instilling positive traits

·LessWrong··

This is the fifth in a series of informal research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. The fourth post can be found here.TLDR: Via adapting the methods of Marks et al and Li et al, we train Gemini 3 Flash to have certain traits/values by midtraining it on documents about how Gemini has those properties, followed by finetuning it on synthetic chat data where it demonstrates those properties. The chat finetuning is effectiv...

Read full article →

Related Articles

Show HN: An e-ink frame that hears birds and draws them as 1800s illustrations
arnemunthekaas · Hacker News · 6h ago
Show HN: Capsule – Single-file web apps that save their data into SQLite
bashtian · Hacker News · 5h ago
How much of F-Droid is LLM generated?
_ZeD_ · Hacker News · 9h ago
Ubuntu 26.10 completes transition to Rust-based coreutils
theanonymousone · Hacker News · 1d ago
America's Driver's License Breach Is a National Security Disaster
hn_acker · Hacker News · 3h ago