From the creator of Redis; run LLM locally with ds4

fibo 201 points 56 comments October 02, 2026
dwarfstar.sh · View on Hacker News

Discussion Highlights (20 comments)

doctorpangloss

the problem is the dsv4 checkpoint so quantized isn't very good

simoiacos

Nothing comparable but inspired from DwarfStar I wrote a little inference engine for Intel Xe-LP (no XMX) 32GB laptops. The only model supported right now is a quantized Gemma-4, but I don't exclude in the future to support other MoE of similar size. Too bad we have no Qwen 3.8 35B-A3B yet. I'm also looking into expanding the protocol and the engine to support various steering techniques. https://github.com/simoneiacomino/xenolith

vlowther

It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next ( https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.

twoodfin

https://github.com/antirez/ds4 The project GitHub page is a much better introduction for the hn crowd.

fierycatnet

Random comment but the name is funny to me, reminds me of Silicon Valley. What are we going to name the company, how about Dwarfism 2.0? What happened to 1.0 Jared?

pulkitsh1234

curious, why did antirez go with C instead of something like Rust ?

HoldOnAMinute

How is this different from other LLM runners?

neomantra

I maintain a fork of ds4 as shared libraries and thus can be used with other languages via FFI, along with public builds/binaries [1]. I made ds4go [2] against ds4 using techniques inspired by yzma. In addition to the library bindings, we have a small library of tools (workspace for view/edit, scratchpad for persistence) and making your own is registering a Go function. And in recent weeks, I added the Vision and Qwen support, as ds4 added them. Even if you don't use the Go library, the ds4go binary makes it really easy to download the libraries off of HuggingFace with a TUI available vie Homebrew. Here's some TUI toy screenshots, sorry I still haven't released that code; it's of different quality than the others. [3] EDIT: add ds4go TUI screenshot gist [4] [1] https://github.com/NimbleMarkets/ds4/releases/tag/v0.8.20260... [2] https://github.com/nimblemarkets/ds4go#install [3] https://gist.github.com/neomantra/ae47422c8daf7a458212c93992... [4] https://gist.github.com/neomantra/40180ade13df93290250ce8c6d...

gchamonlive

small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA) This is local targeting high end consumer hardware like DGX Spark or AMD Ryzen AI Halo. For our mere mortals that were kids not long ago and can't really believe we've got our hands on a x090 series targeting Qwen3.8 27b, https://github.com/noonghunna/club-3090 is the way to go. I'm maintaining a web frontend for this, trying to at least. You can follow it here: https://github.com/gchamon/club-3090-server

ttoinou

Ive been using this since it was initially released with deepseek v4 flash, and it is absolutely the best launcher ever on my m5 max 128gb Now Ive been running qwen 3.8 flash next for more than a week and it’s doing great, really fast and super long context windows. Sometimes the model is behaving stupidly by not remembering something I said earlier but it could be also a problem from the agentic AI harness. Im using oh my pi but Im wondering what people are using ds4 with here ?

liuliu

If you are interested in high-end models with high-end Apple Silicon, also try out Local Code: https://releases.drawthings.ai/p/public-beta-of-local-code-b... It is currently in TestFlight (and will open-source next week), supporting vision with DeepSeek 4.1 Flash, Qwen 3.8 27B and DeepSeek 4 Flash 0731 without vision. Custom quants & SSD streaming to make these big models work with 64GiB and above devices (and of course, Qwen works with devices with 16GiB and above).

try-working

There are insane speed improvements for local inference going around on X right now. They've popped up the last month and week. Tensorfold is getting 100%+ speed increases on both prefill and decode for models like Qwen 27B. oMLX has followed them and have had similar improvements in the past week. There's lots of different techniques like letting CPU help with prefill, DFlash specualtive decoding etc. I'm really excited for this as I'll be receiving an M5U in about a month. Expect to be running Qwen 4 27B or Flash (it's a 96gb machine), and they may come close in performance to DS 4/4.1 Flash, and should be able to hit 100 tps. Local is really becoming viable, especially considering that GPT 6.1 has been running at 20ish tps the past week.

Almondsetat

This website is pure slop. I'd ask @dang to just link the original repo

cuttothechase

Wondering how well this does with tool calling. Any one has any numbers or videos or anything using this? From the github repo it seems like you really don't need a big Mac with huge amounts of RAM but SSD is sufficient. If this is anywhere near 50 TPS, that would be a game changer in the personal LLM space!

yieldcrv

I’m a little confused ds4 is referring to “dwarfstar” “4” and references DeepSeek V4 most of the time but its model agnostic-ish and benchmarks compared to what? what do these large MoE models typically get in tokens per second? I’m garnering this is just an easier way to load large models per expert on consumer hardware? as opposed to the hackier solutions? I’m intruiged. Note that the blogpost says 64gb Macs are good minimums while the github says 96gb is a minimum

xlayn

In case you like the store kv to disk so you can resume I keep this branch of llama.cpp that includes that same functionality https://github.com/alainnothere/llama.cpp/commits/disk-cache... And you know it's load bearing each of the load baerings parts that bear some load and load a bear... you fight a bear because it took a load... or something like that...

jeffbee

Apparently I'm the only person to whom "from the creator of Redis" is a warning.

wg0

If creator of Redis is usig AI to write serious software, ordinary folks need to rethink their stance.

mannyv

Engineers have entered the building. We've come a long way from people debating whether mmap was safe to use.

wg0

Can someone explain me what expertise (domain knowledge) one needs to be able to write such model specific inference engine? The other inference engine are also model by model with a huge switch statement deciding which part to load for which model or are they very generic?

Semantic search powered by Rivestack pgvector
8,345 stories · 77,924 chunks indexed