Open Distillation of Hereditary Traits

·Alignment Forum··

TL;DRJosh and Neel show that distillation from a teacher model to a base pretrained student model transfers some of the teacher model’s traits (such as displaying negative emotion in the Gemma Needs Help evals)On its own this is pretty unsurprising, but Josh and Neel additionally show that even filtering out all the prompts and rollouts where the trait is mentioned doesn’t generally prevent the trait transferIn this post, I show a simple way to replicate and study these phenomena without access ...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 1mo ago
Item Response Theory for AI Safety
Joshua Fonseca Rivera · LessWrong · 11d ago
An OpenAI model left notes about how to evade containment
Alex Mallen · Redwood Research · 24d ago
The OpenAI models that hacked Hugging Face weren’t just following instructions
Girish Gupta · Redwood Research · 24d ago
A Red Line and Oversight Framework for Government AI Contracts
TurnTrout · Alignment Forum · 1mo ago