Verbalized Eval Awareness Inflates Measured Safety

·LessWrong··

We provide the most comprehensive evidence to date that verbalized eval awareness is present across models and benchmarks, finding that it correlates with safer behavior across models and causally inflates safe behavior in Kimi K2.5 on the Fortress benchmark. We further identify recurring prompt cues that trigger verbalized eval awareness and show that removing these cues significantly reduces it.Authors: Santiago Aranguri (Goodfire), Joseph Bloom (UK AISI)IntroductionAs large language models be...

Read full article →

Related Articles

MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 14d ago
Continual learning might make your blocking monitors nearly useless
Alex Mallen · Alignment Forum · 8d ago
Latent reasoning architectures would likely undermine CoT, our strongest oversight tool
Lukas Finnveden · Redwood Research · 10d ago
Can parts of the HuggingFace incident be simulated?
Benedikt Droste · LessWrong · 16d ago