Turning GLM-5.3-Flash into a Jev-like decision model

flxflx 73 points 28 comments September 26, 2026
www.privatemode.ai · View on Hacker News

We found an approach to get Jev-like properties from standard LLMs like GLM-5.3-Flash. The core idea is to craft the input prompt so that the first output token answers the question. This makes it possible to get a decision with a single forward pass. In the blog post, we describe the approach in detail for GLM-5.3-Flash and vLLM. We benchmark this setup against Jev and Laya. We find that our setup is on-par with Jev in terms of accuracy and speed and that it substantially outperforms Laya. Still, in terms of costs per decision, Jev is several x better than our setup. In turn, our setup supports vision inputs.

Discussion Highlights (7 comments)

m4y0u

My question is why not use Jev instead? It's faster and cheaper.

ricardobeat

Everyone is doing this to emulate Jev, but... I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.

ttoinou

Isnt this obvious ? I would have thought people would try such things before deciding they need something like Jev

janalsncm

If you are using an autoregressive decoder (which glm is) it is not “jev-like”. You lose all of the speed advantages that Jev has.

Jabrov

Is this a joke? “Jev-like” properties? People have been using LLMs as classifiers or rankers in a similar way for ages. I feel like we’re losing our minds

prjkt

how is Jev cheaper if I can run locally. 0.5% prefill, 0.1% decode, 99.4% cached, latency is <20ms

walrus01

You can turn any sufficiently smart LLM into yes/no decision model or equivalent. I already have an existing workflow with a two paragraph detailed prompt, that sends pages of stuff to an LLM and asks it to return only 7 JSON objects. Several of those objects are binary "yes or no" choices of like, whether the content contains certain things. You can even do it with small not particularly hard to host local LLMs like a variant of Qwen 3.6 35B A3B or 3.8 27B.

Semantic search powered by Rivestack pgvector
7,763 stories · 72,001 chunks indexed