Hetzner is working on LLM Inference
jonas_scholz
146 points
78 comments
July 24, 2026
Related Discussions
Found 5 related stories in 1159.6ms across 14,736 title embeddings via pgvector HNSW
- UK sovereign LLM inference benjamintnorris · 104 pts · May 15, 2026 · 60% similar
- Show HN: Kelet – Root Cause Analysis agent for your LLM apps almogbaku · 42 pts · April 14, 2026 · 54% similar
- Surpassing vLLM with a Generated Inference Stack lukebechtel · 31 pts · March 10, 2026 · 54% similar
- Show HN: Tiny-vLLM – high performance LLM inference engine in C++ and CUDA yu3zhou4 · 122 pts · May 29, 2026 · 53% similar
- Real-time LLM Inference on Standard GPUs: 3k tokens/s per request NicoConstant · 202 pts · May 29, 2026 · 53% similar
Discussion Highlights (12 comments)
embedding-shape
> The enable_thinking option is worth mentioning. Without it, the model can spend a surprising amount of the completion budget reasoning before it returns a visible answer. Straight up the opposite, which the name makes abundantly clear, with the option it does reasoning, without it it doesn't...
ano-ther
Good to see more developments in this space. I quite like this service, which is a little further than Hetzner and has several models to choose from: https://www.infomaniak.com/en/hosting/ai-services
swiftcoder
It would certainly be interesting to have a highly respected EU-native inference provider, if only to make the regulatory gods happy
mark_l_watson
This seems like a smart move, given their ability to host efficiently. I approve of efforts to make the cost of inference for smaller useful models slowly approach 'close to zero' and there are many good paths for getting there. It is useful for companies to get fast hosting for the class of smaller models they may end up hosting in house.
nubg
Potentially interesting article ruined by AI slop hallucinations like > For now, the API is fast, free, and fun to try. The next hardware announcement will tell us much more than another small model would.
rebelde
Hetzner is very efficient hosting servers Will this be the new division of labor? Americans - best proprietary models Chinese - best open weight models Europeans - best / most efficient inference service
Havoc
Interesting. I could see them perhaps coming in competitive for models that fit into single cards? Less so playing in the big model serving league...climbing into that esp right now would be madness
perelin
Whats definitely missing: a solid (non Mistral) GDPR compliant coding plan / subscription. All offerings are either US or China based. With the newest open weights models this became really interesting imo.
_pdp_
I mean yah... host glm and kimi and I am game.
NetOpWibby
This is interesting because I thought Hetzner was anti-crypto? LLMs aren't the same but they're often lumped in with crypto as "things no one wants."
danlitt
Excited to see the price of every other product they offer triple for no reason.
pmg1991
Very soon we will be having reseller programs for inference, this will be just like web hosting reseller. After big players, small players will also start entering in this field. I'm waiting for that day so that inference will be affordable just like web hosting. 200$ per month is in no way affordable by everyone.