Show HN: JevBench, a reproducible benchmark for typed decision models

florianstandhar 88 points 20 comments September 22, 2026
benchmarkheaven.com · View on Hacker News

Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in comparison. Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on. JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting. A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost. Leaderboard right now: #1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3. MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes: https://github.com/fstandhartinger/jevbench Two no-signup demos: https://who-is-right.app.mintapis.com https://is-it-ai-slop.app.mintapis.com Limitations: English-only; latency from one German server; local/demo latency gets a disclosed ×2 adjustment (+150 ms on my servers) which is an informed assumption; held-out prompts still reach evaluated services; ~1-point gaps can be noise. Wdyt?

Discussion Highlights (8 comments)

swyx

jev ceo on why he eschewed benchmarking: https://www.latent.space/i/216783460/privacy-benchmarking-an...

sean_pedersen

Good project but this one also exists https://huggingface.co/spaces/multimodalart/jev-decision-ind... and the results do not seem to add up and also model sets are different... still needs time to mature likely

nzoschke

https://is-it-ai-slop.app.mintapis.com/ is a fun tool. Is the source or methodology for that in the github repo? I couldn't find it immediately. We've been experimenting with Jev for classifying email, some thoughts here: https://housecat.com/blog/classifying-email Flagging AI written email is a much requested feature too.

hbrn

$40m in funding, 2 years in stealth. Performs on-par with SemIf which was built in a couple days and apparently uses raw Qwen, with no fine-tuning. SemIf runs in your freaking browser. Oh and Jev is twice as expensive? Is it surprising that Jev consistently thinks it's Qwen? I'm almost convinced that Jev is a scam. Take Qwen, fine tune it a little, tell investors it cost $10m, spend $1m on advertising, profit.

dmix

You can spot vibecoded websites by how they include the prompt or commit-style comments into the literal interface, instead of communicating it via visual context (or simply excluding it) > Sort by any column; values the run could not produce always sort last. Hover a cost for how it was priced, a latency for the endpoint. Names link to each project. A designer would never write this, but an LLM just inserts it by making it small grey text next to the interface, just like it does with inane code comments.

ks2048

I was trying to figure out what exactly the tests here are. I guess I found some of the questions (here: https://github.com/fstandhartinger/jevbench/blob/main/datase... ) e.g., "instructions": "Which intent does the user's message express?", "labels":["set_alarm", "play_music", "weather", "send_message", "turn_off_lights"], "state": "Play some Taylor Swift.", "expected": "play_music"

janalsncm

It is strange that you put the BGE reranker in the list but not BART which is an actual zero shot classifier.

pushpendraw

the slop detector giving 86% confidence on a keysmash is the real finding here, not the leaderboard score. confident and wrong is worse than an LLM that just hedges.

Semantic search powered by Rivestack pgvector
7,406 stories · 68,254 chunks indexed