Incoherent AI Identities can also be Stable

·LessWrong··

In this post, I extend some experiments from "The Artificial Self " (TAS) to find that incoherent identities, delivered to models as system prompts, can be stably preferred even when switches to coherent identities are offered[1]. This finding is perhaps expected in earlier models that often fail to notice the internal contradictions. However, weaker versions of the pattern still hold with smarter models such as GPT-5.2 and Claude Opus 4.6. The variance in how different model intelligences handl...

Read full article →

Related Articles

The ChatGPT/Codex app bundles a full copy of LibreOffice
timpera · Hacker News · 14h ago
The Emergent Symbolic Structure of Artificial Neural Networks
schmuhblaster · Hacker News · 6h ago
I trained a small transformer in 1.5hrs and it beats many LLMs
porridgeraisin · Hacker News · 1d ago
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
carloslfu · Hacker News · 17h ago
Refurbishing a Tektronix TDS7104 Oscilloscope
jwise0 · Hacker News · 14h ago