Core Assumptions of ELK

·LessWrong··

In this post I do a brief analysis of the core assumptions of ELK.ContextThe Eliciting Latent Knowledge problem, for the unfamiliar:Suppose we train a model to predict what the future will look like according to cameras and other sensors. We then use planning algorithms to find a sequence of actions that lead to predicted futures that look good to us.But some action sequences could tamper with the cameras so they show happy humans regardless of what’s really happening. More generally, some futur...

Read full article →

Related Articles

The ChatGPT/Codex app bundles a full copy of LibreOffice
timpera · Hacker News · 15h ago
The Emergent Symbolic Structure of Artificial Neural Networks
schmuhblaster · Hacker News · 7h ago
I trained a small transformer in 1.5hrs and it beats many LLMs
porridgeraisin · Hacker News · 1d ago
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
carloslfu · Hacker News · 18h ago
Refurbishing a Tektronix TDS7104 Oscilloscope
jwise0 · Hacker News · 15h ago