LLMs have the capacity for self-imposed steganography

·LessWrong··

OverviewThis is my writeup for my BlueDot Impact - Technical AI Safety Project. In this project I aimed to demonstrate that there is capacity for LLMs to take on steganography capabilities.In terms of AI safety, steganography is of particular interest as it may be used by a misaligned LLM to evade the monitoring of a control protocol. While the main concern is whether these capabilities could naturally occur as a result of high-compute RL pipelines, to make this initial research project tractabl...

Read full article →

Related Articles

LG smart TVs caught logging audio with screen off and snooping on local devices
chris_overseas · Hacker News · 14h ago
Smartphone makers don't bother to comply with EU repairability requirements
mdp2021 · Hacker News · 9h ago
bzip3
tosh · Hacker News · 7h ago
Asahi Linux on M3
mdp2021 · Hacker News · 1d ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 1d ago