DFlash 2: Keep Drafting Parallel
mike-the-brain
83 points
12 comments
August 19, 2026
Related Discussions
Found 5 related stories in 41.2ms across 4,128 title embeddings via pgvector HNSW
- Flux 3 ThouYS · 554 pts · July 24, 2026 · 46% similar
- LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacB Alephinitesimal · 15 pts · August 21, 2026 · 45% similar
- AI at Home Part 2: Multi-GPU Drifting timmmmmmay · 24 pts · August 20, 2026 · 44% similar
- Flux 3 X Mimic: The Next Generation of Video-Action Models kensai · 313 pts · July 24, 2026 · 42% similar
- The New AI Superpowers: Focus and Followthrough mooreds · 184 pts · July 26, 2026 · 42% similar
Discussion Highlights (5 comments)
verdverm
vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816
adefa
I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.
hypfer
Amazing tech > An agent writes in an afternoon what a chatbot writes in a month But can you just.. not. Your tech is so good, it speaks for itself. Don't ruin that.
sarjann
Great news, has made low memory bandwidth model usage so much nicer.
ilc
Watch the video carefully. DFlash2's tool call fails on python syntax. Usually models in this class nail things like that 1 shot, which the other side did. I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.