Eliciting hidden knowledge from monitors with NLAs

·LessWrong··

Aleksandr Bowkis* and David Africa*TL;DRChain of thought (CoT) monitorability may be fragile, and natural language autoencoders (NLAs) may provide a helpful, decorrelated monitoring surface.We tried to read NLAs from the monitor itself, where the NLA readout surfaces what the monitor internally represents while judging an agent's trajectory.NLAs can be useful for monitoring in two ways:Monitor-side: Eliciting latent capabilities from weak monitors by surfacing unverbalised knowledge of reward ha...

Read full article →

Related Articles

Field measurements of neighborhood-scale air temperature impacts of data centers
cwwc · Hacker News · 15h ago
Solo – a .so loader for static Linux binaries
zX41ZdbW · Hacker News · 9h ago
Linux 7.3 improves performance when running out of vRAM
flaburgan · Hacker News · 1d ago
Memory prices climb 500% in 12 months
haunter · Hacker News · 1d ago
A 3D fruit fly on macOS desktop powered by the real FlyWire connectome
phoenix120 · Hacker News · 11h ago