Measuring Activation Control in LLMs

·LessWrong··

TL;DRInspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and potentially even influencing their own training.We ran this on 25 open weight models ranging from 4B to 744B. We find that most language models are able to not only increase...

Read full article →

Related Articles

DeepSeek V4 Pro 0813
explosion-s · Hacker News · 12h ago
Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot
gavinhking · Hacker News · 14h ago
Tailscale Traces Database Corruption to 16y/o SQLite WAL-Reset Bug
ropbear · Hacker News · 14h ago
What sort of maths are LLMs good at?
ColinWright · Hacker News · 18h ago
England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 1d ago