LLMs have the capacity for self-imposed steganography

·LessWrong··

OverviewThis is my writeup for my BlueDot Impact - Technical AI Safety Project. In this project I aimed to demonstrate that there is capacity for LLMs to take on steganography capabilities.In terms of AI safety, steganography is of particular interest as it may be used by a misaligned LLM to evade the monitoring of a control protocol. While the main concern is whether these capabilities could naturally occur as a result of high-compute RL pipelines, to make this initial research project tractabl...

Read full article →

Related Articles

ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 1d ago
Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 1d ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 2d ago
DeepSeek Elastic Compute (DSec)
shenli3514 · Hacker News · 17h ago
Ask HN: Who's still keeping a DOS machine up because the business depends on it?
mlaux · Hacker News · 1d ago