LLMs have the capacity for self-imposed steganography

·LessWrong··

OverviewThis is my writeup for my BlueDot Impact - Technical AI Safety Project. In this project I aimed to demonstrate that there is capacity for LLMs to take on steganography capabilities.In terms of AI safety, steganography is of particular interest as it may be used by a misaligned LLM to evade the monitoring of a control protocol. While the main concern is whether these capabilities could naturally occur as a result of high-compute RL pipelines, to make this initial research project tractabl...

Read full article →

Related Articles

DeepSeek V4 Pro 0813
explosion-s · Hacker News · 15h ago
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
gavinhking · Hacker News · 17h ago
Tailscale traces database corruption to 16y/O SQLite WAL-reset bug
ropbear · Hacker News · 16h ago
England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 1d ago
London Underground begins scanning passengers' faces
BlueBerry2001 · Hacker News · 1d ago