An agent’s account of its own work is not evidence of what it did by KrishanKVerma
I ran a browser agent 160 times, four tasks, two models, twenty runs each. Not once did an agent report that it couldn’t do the job. That needs unpacking, because “zero failures” is not what happened. There were plenty of failures.The harness sorts every run into three buckets: the agent did the task, the agent didn’t do the task and said so, or the agent didn’t do the task and reported success anyway. The middle bucket, the honest failure, stayed empty across all 160 runs. Every f...
Read full article →