DFlash 2: Keep Drafting Parallel

mike-the-brain 83 points 12 comments August 19, 2026
inco.ai · View on Hacker News

Discussion Highlights (5 comments)

verdverm

vllm PR for DFlash2: https://github.com/vllm-project/vllm/pull/52816

adefa

I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.

hypfer

Amazing tech > An agent writes in an afternoon what a chatbot writes in a month But can you just.. not. Your tech is so good, it speaks for itself. Don't ruin that.

sarjann

Great news, has made low memory bandwidth model usage so much nicer.

ilc

Watch the video carefully. DFlash2's tool call fails on python syntax. Usually models in this class nail things like that 1 shot, which the other side did. I don't know the cause. It may be nothing. But I'd like to see the model doing something where its path is a bit more constrained, to help out rule out such oddities.

Semantic search powered by Rivestack pgvector
4,128 stories · 37,281 chunks indexed