We should consider how long monitoring is reliable for during RL

·LessWrong··

Epistemic status: I am new to AI Safety and am writing blogs to gain context. This blog post was formed from discussions with Aidan Ewart and Jonathan Bostock, but they do not necessarily endorse this post.TL;DRGiven recent examples of AI misbehaviour during training episodes, AI companies might want to start using monitoring during training as well as deployment. But this might have the effect of training the AIs to simply evade the monitors. Depending on the specifics of the monitoring protoco...

Read full article →

Related Articles

England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 17h ago
Stealing Reasoning Traces from Proprietary LLM APIs
quantumgarbage · Hacker News · 16h ago
CFTC declares market emergency, orders Kalshi to continue to operate in New York
michaefe · Hacker News · 5h ago
London Underground begins scanning passengers' faces
BlueBerry2001 · Hacker News · 20h ago
Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 1d ago