Benchmarking Jev against no-CoT LLMs

·LessWrong··

TL;DR. We ran Jev 1.13 — TypeSafe's non-autoregressive model, which answers questions with probabilities and cannot emit text, on the ThinkFast no-CoT suite and Neel Nanda’s NCRI.Jev’s capabilities are very jagged. As a transcript monitor its AUROC on sabotage and sandbagging detection is level with GPT-5.5 and Opus 4.7 without CoT, at about 1/200th of Opus 4.7’s cost. On multiple-choice questions it ranges from ahead of GPT-5.5 (GPQA-Diamond) to well behind it (Sally-Anne). It does badly on tas...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 1d ago
AnyPS5: Port PS5 binaries to PC without emulation (87% system libraries mapped)
Fe2O3 · Hacker News · 14h ago
JetBrains reported a net financial loss first time in its tracked history
thw_9a83c · Hacker News · 1d ago
OpenTPU – An open-source AI accelerator, developed by AI
fsbonetto · Hacker News · 21h ago
Shipping JPEG XL in Chrome
AshleysBrain · Hacker News · 2h ago