A proposal for a highly effective AI safety org

·LessWrong··

TLDR: an org that pays people to "just read the fucking transcripts"; a large amount of people reading anonymized claude code/RL/eval transcripts flagged by a very high recall low precision monitor could catch warning shots, reward hacking and general weird stuff without needing to absorb any good people and this doesn't seem to exist. $5M/month could pay 1000 people to process literally all tokens in a frontier RL run and would catch ~15 serious incidents per month in bearish estimates."Can't t...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 21h ago
Asahi Linux Now Officially Supports Apple M3 Macs – With Caveats
mdp2021 · Hacker News · 3h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 21h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago