Measuring Activation Control in LLMs

·LessWrong··

TL;DRInspired by the introspective awareness and CoT controllability papers, we made a benchmark to measure how well models can control their activations while completing a simple task. We are motivated by the concern that highly introspective models could control their activations, confounding probes and other monitors, and potentially even influencing their own training.We ran this on 25 open weight models ranging from 4B to 744B. We find that most language models are able to not only increase...

Read full article →

Related Articles

Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 1d ago
ASML says it sold 'absolutely nothing' in Europe in 2026
MC995 · Hacker News · 1d ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 1d ago
DeepSeek Elastic Compute (DSec)
shenli3514 · Hacker News · 12h ago
Ask HN: Who's still keeping a DOS machine up because the business depends on it?
mlaux · Hacker News · 1d ago