Show HN: JevBench, a reproducible benchmark for typed decision models

·Hacker News··

Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in comparison.Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.Leaderboard right now: #1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3. MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes:https://github.com/fstandhartinger/jevbenchTwo no-signup demos:https://who-is-right.app.mintapis.comhttps://is-it-ai-slop.app.mintapis.comLimitations: English-only; latency from one German server; local/demo latency gets a disclosed ×2 adjustment (+150 ms on my servers) which is an informed assumption; held-out prompts still reach evaluated services; ~1-point gaps can be noise.Wdyt?

Read full article →

Related Articles

Claude Opus 5.5
km144 · Hacker News · 5h ago
I asked Meta’s Muse for its filesystem and it sent me 6.8GB
Aeroi · Hacker News · 6h ago
There's a high chance of devices being sold with GrapheneOS preinstalled in 2027
Cider9986 · Hacker News · 4h ago
AMD's random number generator can't generate a 0?
BruceEel · Hacker News · 13h ago
NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 1d ago