Model self-identification could be subliminally transferred

·LessWrong··

Identity questions seem hard to get right. Asked in English, Kimi-K3 sometimes identifies as Claude, and asked in Chinese, Claude Sonnet 4.6 sometimes claims it is DeepSeek. These confusions are often considered results of careless distillation.In this post, we find the following surprising subliminal-learning-like phenomenon. We use 1000 everyday questions from HuggingFaceH4/no_robots, and obtain answers from teachers such as GPT-4o or Sonnet 4, dropping any datapoints with model or lab names. ...

Read full article →

Related Articles

Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 13h ago
JetKVM Mini
taubek · Hacker News · 7h ago
google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 1d ago
Revolut confirms customer data breach through fake government requests
tdrz · Hacker News · 5h ago
After Math
throwaway81523 · Hacker News · 11h ago