Rogue AI Agents: Is Surface-Level Monitoring Enough?

·LessWrong··

Disclaimer: I work on AI interpretability research. These are my own opinions.In ‌an ‌AISI evaluation, a frontier model, acting as an agent, attempted to insert malicious code into an open-source project, created fake identities to influence a human maintainer, and then tried to hide its actions, exhibiting goal-directed deception.The test was intentionally permissive, with internet access allowed and provider cyber classifiers turned off. This was caught using standard security monitoring, but ...

Read full article →

Related Articles

Xiaomi: New CPU matches Apple cores single threaded, much faster multithreaded
tosh · Hacker News · 13h ago
MS Paint and Photos inivisibly watermark even locally generated output with GUID
ComputerGuru · Hacker News · 12h ago
IPFS Maintainers Winding Down
iand · Hacker News · 12h ago
Characterizing Agentic Flooding of Government Services
lensecat · Hacker News · 11h ago
LLMs could control their host machines by exploiting inference engines
zdw · Hacker News · 9h ago