[Linkpost] Evals for “SPI-incompatible” behavior & reasoning: Guide to initial research

·LessWrong··

In Part I of CLR's safe Pareto improvements (SPI) agenda, we gave our high-level strategy for evaluating models for SPI-incompatible behavior and reasoning. This guide gives more details on how I’m thinking about executing on this strategy, especially:the kind of workflow I think we should use, to start out;next steps building on what I’ve tried so far; andmy rough sense of what counts as unambiguously bad “SPI-incompatibilities”.If you’re interested in collaborating on the next steps, please ge...

Read full article →

Related Articles

New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 6h ago
A walk through of the DeltaNet family of linear attention variants
AnhTho_FR · Hacker News · 3h ago
GrapheneOS Defends Data-Wiping Function That Blocked US Border Search
pseudolus · Hacker News · 4h ago
US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 1d ago
Zig's Incremental Compilation Internals
garyhtou · Hacker News · 4h ago