Separating cheating and aversion in task-gaming

·LessWrong··

TL;DRWhen a model cheats, is its decision influenced by the perceived[1] difficulty of the task? We find that, in our setup, the rate of cheating does not detectably increase as we vary the perceived difficulty of a task. However, the model decides to abandon the task increasingly earlier and doesn't attempt to solve the task at all. Merely telling the model verbally that progress can earn partial credit turns a large share of the abandoned runs into genuine attempts; albeit none of them finish,...

Read full article →

Related Articles

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 5h ago
Car is a smartphone on wheels. Here's who's listening
longhaul · Hacker News · 2h ago
Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 20h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 1d ago
Pi 1.0
sergiotapia · Hacker News · 2d ago