Eliciting hidden knowledge from monitors with NLAs

·LessWrong··

Aleksandr Bowkis* and David Africa*TL;DRChain of thought (CoT) monitorability may be fragile, and natural language autoencoders (NLAs) may provide a helpful, decorrelated monitoring surface.We tried to read NLAs from the monitor itself, where the NLA readout surfaces what the monitor internally represents while judging an agent's trajectory.NLAs can be useful for monitoring in two ways:Monitor-side: Eliciting latent capabilities from weak monitors by surfacing unverbalised knowledge of reward ha...

Read full article →

Related Articles

GLM-5.3 is now open-weight
jeudesprits · Hacker News · 1d ago
Samsung's Processing-in-Memory (PIM)
ingve · Hacker News · 10h ago
EPA says power for data centers can sidestep pollution laws
Levitating · Hacker News · 1d ago
Just the rumour of a bug is enough to find an exploit these days
avsm · Hacker News · 1d ago
Does the Sumerian King List Align with Paleoclimate Events?
dev_l1x_be · Hacker News · 17h ago