Classifier Context Rot: Monitor Performance Degrades with Context Length

·LessWrong··

Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely contain transcripts longer than 100K tokens.We show that when used as classifiers, current frontier models fail to notice dangerous actions more often in longer transcripts. In particular, on MonitorBench, Opus 4.6, GPT 5.4, and Gemini 3.1 miss these actions 2x to 30x more often when we prepend 800K tokens of benign act...

Read full article →

Related Articles

Federal judge calls Flock 'indiscriminate mass surveillance'
sbulaev · Hacker News · 6h ago
Kolibri: A Sovereign Open-Weight Model
bastitx · Hacker News · 19h ago
Pi 1.0
sergiotapia · Hacker News · 2d ago
Updates to Full Disk Access in macOS
notfirstpost · Hacker News · 1d ago
FTL: A new operating system for clouds
romac · Hacker News · 13h ago