Z-Lab's DFlash-2 boosts LLM throughput with diffusion-based drafting
Z-Lab and inco.ai released DFlash-2, an upgrade to the DFlash draft-token prediction technique built on a diffusion model. It adds a Lightweight Path Selector for token ordering and Local Convolution to counter suffix decay. Only Qwen3.8-27B and Muse-Glimmer-30B support it so far, with llama.cpp support via PR #27342.
- DFlash-2 adds a Path Selector and Local Convolution to DFlash diffusion drafting
- Available models: Qwen3.8-27B-DFlash2 and Muse-Glimmer-30B-DFlash2 with GGUF builds
- llama.cpp support is only in PR #27342, not yet in master
- Recommended settings: --spec-draft-n-max 4, --temp 1.0, --top-p 0.95
Read next
AI