Benchmarking Jev against no-CoT LLMs
TL;DR. We ran Jev 1.13 — TypeSafe's non-autoregressive model, which answers questions with probabilities and cannot emit text, on the ThinkFast no-CoT suite and Neel Nanda’s NCRI.Jev’s capabilities are very jagged. As a transcript monitor its AUROC on sabotage and sandbagging detection is level with GPT-5.5 and Opus 4.7 without CoT, at about 1/200th of Opus 4.7’s cost. On multiple-choice questions it ranges from ahead of GPT-5.5 (GPQA-Diamond) to well behind it (Sally-Anne). It does badly on tas...
Read full article →