Rogue AI Agents: Is Surface-Level Monitoring Enough?
Disclaimer: I work on AI interpretability research. These are my own opinions.In an AISI evaluation, a frontier model, acting as an agent, attempted to insert malicious code into an open-source project, created fake identities to influence a human maintainer, and then tried to hide its actions, exhibiting goal-directed deception.The test was intentionally permissive, with internet access allowed and provider cyber classifiers turned off. This was caught using standard security monitoring, but ...
Read full article →