Low Probability Estimation in Language Models

·ARC··

ARC recently released our first empirical paper: Estimating the Probabilities of Rare Language Model Outputs. In this work, we construct a simple setting for low probability estimation — single-token argmax sampling in transformers — and use it to compare the performance of various estimation methods. ARC views low probability estimation as a potential technique for mitigating worst-case behavior from AI, including deceptive alignment; see our previous theory post on estimating tail risks in neu...

Read full article →

Related Articles

Reducing synthetic markers makes some SDF false facts linearly indistinguishable from pretraining-acquired knowledge
Jason Zeng · LessWrong · 38m ago
MIT's New Method Flags AI Models Trained on CASM Without Generating It
sdoering · Hacker News · 2mo ago
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
sbulaev · Hacker News · 15d ago
Continual learning might make your blocking monitors nearly useless
Alex Mallen · Alignment Forum · 9d ago
Latent reasoning architectures would likely undermine CoT, our strongest oversight tool
Lukas Finnveden · Redwood Research · 10d ago