The Pragmatic Interpretability Trap

·LessWrong··

TLDR: Pragmatic interp sounds great in the sense that you get to keep interp tools while actually moving safety metrics, but looking a bit closer it's kinda a trap. You pay interp's overhead but get judged against black-box baselines that don't, so the work that survives is whatever cleared that bar, not whatever produced understanding. The two scoreboards (understanding vs intervention) don't actually run side by side; the metric one eats the other, and stuff like NLAs end up failing both, you ...

Read full article →

Related Articles

Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 10h ago
Pi 1.0
sergiotapia · Hacker News · 2d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 4h ago
Court agrees with EFF: Utah's VPN law demands a technical impossibility
hn_acker · Hacker News · 1d ago