Model self-identification could be subliminally transferred

·LessWrong··

Identity questions seem hard to get right. Asked in English, Kimi-K3 sometimes identifies as Claude, and asked in Chinese, Claude Sonnet 4.6 sometimes claims it is DeepSeek. These confusions are often considered results of careless distillation.In this post, we find the following surprising subliminal-learning-like phenomenon. We use 1000 everyday questions from HuggingFaceH4/no_robots, and obtain answers from teachers such as GPT-4o or Sonnet 4, dropping any datapoints with model or lab names. ...

Read full article →

Related Articles

AI's top startups are barely publishing their research
YeGoblynQueenne · Hacker News · 14h ago
Document-borne AI worms can self-propagate through Copilot for Word
Canopy9560 · Hacker News · 1d ago
Keychron announces first open-source firmware for gaming mice
JLO64 · Hacker News · 19h ago
Turning a dumb AC unit smart (without losing my security deposit)
austinallegro · Hacker News · 17h ago
Handbook.md shows that long policy documents do not reliably govern agents
spIrr · Hacker News · 23h ago