A proposal for a highly effective AI safety org

·LessWrong··

TLDR: an org that pays people to "just read the fucking transcripts"; a large amount of people reading anonymized claude code/RL/eval transcripts flagged by a very high recall low precision monitor could catch warning shots, reward hacking and general weird stuff without needing to absorb any good people and this doesn't seem to exist. $5M/month could pay 1000 people to process literally all tokens in a frontier RL run and would catch ~15 serious incidents per month in bearish estimates."Can't t...

Read full article →

Related Articles

Paint.net 5.2 alpha now runs on Linux
judah · Hacker News · 6h ago
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
jakobgreenfeld · Hacker News · 9h ago
GrapheneOS says Pixel 11 has MTE support after all
user_7832 · Hacker News · 9h ago
Aging brains blend memories together instead of just forgetting them
mdp2021 · Hacker News · 10h ago
Biggest dark matter detector spots a single weird particle
randycupertino · Hacker News · 9h ago