Training on probes: Research ideas

·LessWrong··

RecapSequel to Previous Post. This post might not make sense without it.Training on probes might let our judgments on easy domains generalize to harder domains by leveraging an AI's model of the world. Taking the gradient of the probe teaches the model to fool the probe, but doing RL against the probe doesn't teach the model to fool it!Following the brain-like storyMy hardwired instincts for eating fresh fruit have recruited my learned knowledge about supermarkets, and now I want to go to the su...

Read full article →

Related Articles

LG TVs caught spying even when offline or on standby
sbulaev · Hacker News · 3h ago
Navier-Stokes – Tristan Buckmaster [pdf]
procedurecall · Hacker News · 14h ago
DHS 'Predictive Policing' Unit Is Analyzing Americans' Financial Habits
abraham · Hacker News · 5h ago
Google DeepMind Releases AlphaGenome Atlas
utiiiD · Hacker News · 4h ago
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
stared · Hacker News · 4h ago