Evaluating Offline Monitoring of Internal AI Agents

·LessWrong··

This work was conducted during the GovAI Winter Fellowship 2026. Full reportExecutive SummaryFrontier AI companies use offline monitoring to address risks from internally deployed AI agents. AI developers increasingly rely on AI agents for internal work, including for safety research and model training. At the same time, these companies are concerned that a misaligned model could exploit this access to take concerning actions, such as sabotaging efforts to understand the risks posed by AI. To id...

Read full article →

Related Articles

England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 1d ago
London Underground begins scanning passengers' faces
BlueBerry2001 · Hacker News · 1d ago
Facebook ads are so hard to block that uBlock Origin stopped filtering them
bundie · Hacker News · 6h ago
Stealing Reasoning Traces from Proprietary LLM APIs
quantumgarbage · Hacker News · 1d ago
Beef and dairy drive 41% of biodiversity damage linked to global farmland
robtherobber · Hacker News · 9h ago