We should consider how long monitoring is reliable for during RL

·LessWrong··

Epistemic status: I am new to AI Safety and am writing blogs to gain context. This blog post was formed from discussions with Aidan Ewart and Jonathan Bostock, but they do not necessarily endorse this post.TL;DRGiven recent examples of AI misbehaviour during training episodes, AI companies might want to start using monitoring during training as well as deployment. But this might have the effect of training the AIs to simply evade the monitors. Depending on the specifics of the monitoring protoco...

Read full article →

Related Articles

Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 12h ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 1d ago
Excel now supports multiple values in a single cell
luispa · Hacker News · 12h ago
Toyota is taking the Corolla electric
cisc · Hacker News · 2d ago
Ask HN: Who's still keeping a DOS machine up because the business depends on it?
mlaux · Hacker News · 13h ago