Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms
firelex
425 points
157 comments
September 28, 2026
Related Discussions
Found 5 related stories in 81.0ms across 7,945 title embeddings via pgvector HNSW
- Jev vs. Kev: open-source Jev alternative tested side by side felix089 · 12 pts · September 25, 2026 · 62% similar
- Show HN: JevBench, a reproducible benchmark for typed decision models florianstandhar · 88 pts · September 22, 2026 · 61% similar
- Kev: Tiny Jev-like family of decision models built on top of Qwen3.5 tosh · 426 pts · September 21, 2026 · 60% similar
- Turning GLM-5.3-Flash into a Jev-like decision model flxflx · 73 pts · September 26, 2026 · 59% similar
- Introducing System One Models and Jev albelfio · 1080 pts · September 15, 2026 · 59% similar
Discussion Highlights (20 comments)
firelex
Hi HN. Jeff is a set of small, open-weight Qwen3.5 and Gemma fine-tunes for zero-shot classification, with respectable out-of-the-box performance, meant to be slotted right into code (or fine-tuned further as needed). You give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, with no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in about 28 ms on an M4 Max. Apache 2.0, with a Jev-compatible API (I'm not affiliated with TypeSafe). When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware. So everything ran at home: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing, all monitored from my phone over Tailscale. The caveat: the published Jev and AutoJev numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86-89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64-68% against Jev's 94%, and about 50% on JevBench's hard tier against 73%. That isn't surprising, and I don't think it matters: no 0.8B or 2B model reasons like a large one, and nobody should expect it to. These are extremely fast judgement-callers. In one of my apps I use the 0.8B for voice navigation; a quick fine-tune (about half an hour on one GPU) took it from 32% to 96% on held-out commands, at about 40 ms per decision. The fun part: games, as a zero-shot test. Games aren't the ideal zero-shot test, but they're fun, and TypeSafe did it with Jev too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one; the options say what each move leads to, never which one is right. Over 20 episodes each: - Doom (ViZDoom): Jeff 0.8B 6.55 kills per episode, the same as a hand-coded bot and as Jev's published run. Jev's prompt spells out the aiming rule and takes about 212 ms per call over its API; Jeff gets "the nearest monster is a little to your left" and decides in about 29 ms on my Mac. - Frogger: 10.3 crossings, level with the hand-coded bot (10.25), and 10x the untrained base model (1.0). - Pac-Man: 57 of 98 pellets, about 60% of the bot's score and 2x the untrained model. Videos of every run are linked in the README. Lessons learned: - System 1 models are here to stay. Being able to process unstructured data at software speed inside an app is extremely powerful, and being able to do it locally is fantastic. - A small model is a classifier, not a planner. Models of 0.8B-2B don't reason like Qwen3.8-27B or Jev, and they don't need to: present the options the right way and you get 40+ decisions per second, depending on your hardware. - Fine-tune it if needed. If zero-shot isn't good enough for your task, a short fine-tune on your own examples is. - Wording matters enormously. Giving Frogger's final step the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Before that, the frog just stayed on the last log. - Bigger isn't better. The untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers turning away from the nearest monster), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B. - Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, but not reliably, and in a real-time loop the mistakes compound.
danbrooks
This type of project looks extremely useful. There was a lot of buzz around Jev, but having models that run locally and can be fine-tuned is extremely helpful.
AgentMasterRace
I compared it to Jev in my current use cases and it's very inaccurate. 70% vs 94% . for classification, it's unacceptable.
k__
Von 1.2 had a better Doom score :D https://github.com/wfzyx/von
adrithmetiqa
Forgive my lack of understanding but how long before Jev type functionality is just built straight into all frontier models?
bilekas
Can we get a price comparisson ? Edit: Running them for the masses.
velominati
Typesafe has been quite about the underlying technology behind Jev. Given the speed and cost my hypothesis is that it doesn’t input tokens the way that LLMs do, ie iterating over every word and drawing the connections between each. That is an o(n^2) problem which is why LLMs are so expensive as they scale.
trebligdivad
What proportion of commercial LLM use is classification? I'm just wondering what happens to business AI spending/data centre usage when they realise they don't need full LLMs.
yesthisiswes
My name jeff
drzhouq
Anything like this in the VLM side? Classification on images...
ijustlovemath
Isn't jev just a less nuanced classifier? What am I missing?
zeroCalories
I've always been more afraid of these types of models than LLMs. These are what enable mass surveillance at scale and autonomous real time combat drones. Now they are spreading and being optimized. Gg.
phlipski
For those of a certain age - the fact that Askjev.com is still available astounds me.
qtalen
Awesome, I was just looking for a decision-making model that can be deployed locally, and here you are. Thanks a lot!
desireco42
I think Jev is still the king and while I really appreciate this and other similar projects, Jev gives assurance and value that is hard to beat. Well.. this works locally which is always best, even if slower.
imranq
Everyone saying you could replace Jev or decision type models with an LLM with bolted schema output constraining are missing the point completely. Its about extreme speed and cost effectiveness with high quality, neither of which you are going to get with LLMs even with these KV-cache tricks
folayii
83.1 vs 83.0 on your panel against 70 vs 94 in someone's actual use case is the whole story with zero-shot classification. Any sense of what the 0.8B does on a phone NPU instead of an M4 Max? That's what decides if it's shippable on device.
Kvarnek
Training 0.8B models at home with that latency is seriously impressive. What kind of hardware setup did you use for training?
olwmc
Sorry, do we have actual clear implementation details for Jev? I keep seeing these "recreations" or "Do Jev at home" but do we have access to their architecture? I haven't even used the product, I just find it strange.
hadlock
If you're looking for a more impressive doom example, I put together laya-duum which uses open source micropython implementation of doom (duum) and freeware freedoom1.wad. It uses the standard jev api and will play through the first two levels to completion: https://github.com/Hadlock/laya-duum