Jeeves. Reasoning improves Jev-like decision models
nicowaltz
234 points
89 comments
September 29, 2026
Related Discussions
Found 5 related stories in 99.6ms across 8,041 title embeddings via pgvector HNSW
- Jev vs. Kev: open-source Jev alternative tested side by side felix089 · 12 pts · September 25, 2026 · 65% similar
- Show HN: JevBench, a reproducible benchmark for typed decision models florianstandhar · 88 pts · September 22, 2026 · 65% similar
- Jev in practice: typed decisions, scoped authority niyikiza · 28 pts · September 23, 2026 · 64% similar
- Kev: Tiny Jev-like family of decision models built on top of Qwen3.5 tosh · 426 pts · September 21, 2026 · 62% similar
- Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms firelex · 425 pts · September 28, 2026 · 62% similar
Discussion Highlights (20 comments)
alienbaby
Just curious, where has this term 'noul' come from for yes/no ansers? /a bit more digging and.. A Noul performs a Bernoulli trial—an experiment with exactly two outcomes (yes or no)—but instead of picking one, it returns the calibrated probability (ranging from 0.0 to 1.0) that the statement is true. I hate it :)
hjun1052
If the model does autoregressive reasoning before the decision, doesn't that give up much of what a Jev-style model buys you (a single forward pass, cheap calibrated probabilities)? Or is the point mainly to keep the typed output and probability interface while getting better accuracy on harder cases?
zerop
Are there "good" Open source Decision models built on Gemma-4 and also trainiable on own data?
raverbashing
Jeeves, that's a name I haven't heard in a long time...
woadwarrior01
This isn't really surprising. LLM reasoning and before that, chain of thought prompting are essentially forms of test-time compute scaling.
swader999
Seems like this is the way, a hybrid approach where some of the pipeline will be jev like and some traditional LLM depending on the nature of the work.
phplovesong
So "askjeeves" has been resurrected?
thm
Ask Jeeves - Only took us 30 years to come full circle.
RamblingCTO
Super dope. If it would ship as prod ready code supporting mps as well that would be even doper. But funny that jev is getting its lunch eaten apparently in under two weeks?
mxkuzn
interesting bench list, what about benchmark against smaller or bigger models? 9B looks too huge for small like laya, and too small for llm-level decisions.
captainbland
See if it can beat Jev's Pokémon benchmark
Naitik88
what about benchmark against smaller or bigger models? 9B looks too small for llm-level decisions.
AnodicElegy
I'm surprised we haven't seen a "Jehovah" yet.
sharih
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
esafak
Jev-like models give calibrated decision probabilities, but at low accuracy. So why didn't they show both??
TN1ck
I just did a run with a benchmark I just used to test other models against. (It's about detecting irony in german soccer tweets). On my M5 Pro with 48GB it took over 30min to decide on just 100 tweets, the thinking definitely takes long. It performed quite below Jev, but above other open decision models I tested (68 correct vs 79 correct for Jev - see [1]). I'm running it for the moderation benchmark as well, but that will probably take a few hours on my machine. [1] https://tn1ck.com/blog/jevdit
loclol101
How general really are these jev type models? Has anyone done any broad very cross-domain eval on them?
swingboy
Any good classifiers like this or Jev that support image input?
quantized_state
The diffusion drafter adaptation is nice
theanonymousone
This reminds me of "on-premise cloud".