An agent’s account of its own work is not evidence of what it did by KrishanKVerma

·Nuno Sempere··

I ran a browser agent 160 times, four tasks, two mod­els, twenty runs each. Not once did an agent re­port that it couldn’t do the job. That needs un­pack­ing, be­cause “zero failures” is not what hap­pened. There were plenty of failures.The har­ness sorts ev­ery run into three buck­ets: the agent did the task, the agent didn’t do the task and said so, or the agent didn’t do the task and re­ported suc­cess any­way. The mid­dle bucket, the hon­est failure, stayed empty across all 160 runs. Every f...

Read full article →

Related Articles

GPT 7 (OpenAI) release date
Bayesian · Manifold Markets · 17h ago
Research Report: A genetic basis for potential nociception in cochineal bugs (Dactylopius coccus; Hemiptera: Dactylopiidae) by Meghan Barrett
Meghan Barrett · Nuno Sempere · 18h ago
We know who lives inside wild fish. We don’t know what happens in there by Pablo Sar
Pablo Sar · Nuno Sempere · 2d ago
Will Grok 4.7 be released on 9/11?
Eli Goldfine · Manifold Markets · 3d ago
Will I solve an unsolved math problem with AI in September?
Bayesian · Manifold Markets · 4d ago