Monitoring computer-use agents in OSWorld-control

Chris Harig·LessWrong·Community·June 1, 2026

TL;DR: I built a control setting for computer-use agents and studied a few monitoring setups. I found that attackers can leverage the opacity of computer-use tool calls and gaps between screenshot-action pairs to access credentials or exfiltrate data to other devices on a network without being caught by a monitor. These results validate a possible attack strategy but are not comprehensive evidence that the strategy is uniquely dangerous, given the small sample sizes and limited task set.The sett...

Read full article →

Monitoring computer-use agents in OSWorld-control

Related Articles