DFlash 2: Keep Drafting Parallel
mike-the-brain
83 points
12 comments
August 19, 2026
Related Discussions
Found 5 related stories in 106.8ms across 8,687 title embeddings via pgvector HNSW
- Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency _ache_ · 11 pts · August 26, 2026 · 46% similar
- Flux 3 ThouYS · 554 pts · July 24, 2026 · 46% similar
- DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression mfiguiere · 94 pts · September 17, 2026 · 46% similar
- Qwen3.8-Flash-Next tosh · 657 pts · August 26, 2026 · 45% similar
- LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacB Alephinitesimal · 15 pts · August 21, 2026 · 45% similar
Discussion Highlights (5 comments)
verdverm
vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816
adefa
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
hypfer
Amazing tech > An agent writes in an afternoon what a chatbot writes in a month But can you just.. not. Your tech is so good, it speaks for itself. Don't ruin that.
sarjann
Great news, has made low memory bandwidth model usage so much nicer.
ilc
Watch the video carefully. DFlash2's tool call fails on python syntax. Usually models in this class nail things like that 1 shot, which the other side did. I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.