Inoculate or Reflect? Two training interventions under prompting, steering, and patching

·LessWrong··

Anthropic's recent paper, Verbalizable Representations Form a Global Workspace in Language Models, contains a small experiment near the end that we found more interesting than the main findings.The technique is called Counterfactual Reflection Training (CRT). The model is fed a partial transcript in its context window, followed by an interruption with a question about what matters in that situation, and is trained only on its answer to that question. It is never trained on a corrected action in ...

Read full article →

Related Articles

Clinical failure rates over the decades: yikes
EA-3167 · Hacker News · 20h ago
GM Backs Sodium Ion Batteries for U.S. Grid Storage
rbanffy · Hacker News · 22h ago
Producing ammonia and fertiliser using wind power in Morris, Minnesota
gritzko · Hacker News · 1d ago
ARC-AGI Leaderboard
rzk · Hacker News · 1d ago
Running a 28.9M parameter LLM on an $8 microcontroller
boveyking · Hacker News · 1d ago