Qualia-based emotion steering makes llms attribute conscious states to themselves

·LessWrong··

Summary results/takeawaysI steer Qwen3-32B and 235B along qualia-related emotion directions (blissful, tormented, terrified, serene, etc.) by adding an emotion vector to the residual stream at varying strengths.[1] Then, through a series of forced-choice YES/NO questions, I find that each model becomes much more likely to claim that it is conscious, that it feels and wants things, that it has introspective access to its inner states, and that those states matter morally. Since concept injection ...

Read full article →

Related Articles

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 2h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 17h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 1d ago
Pi 1.0
sergiotapia · Hacker News · 2d ago
The work by Valve's Timur Kristóf on improving old AMD GPUs on Linux
speckx · Hacker News · 20h ago