chiprook
← AI
AIOctober 9, 2026, 12:49

Z-Lab's DFlash-2 boosts LLM throughput with diffusion-based drafting

Z-Lab and inco.ai released DFlash-2, an upgrade to the DFlash draft-token prediction technique built on a diffusion model. It adds a Lightweight Path Selector for token ordering and Local Convolution to counter suffix decay. Only Qwen3.8-27B and Muse-Glimmer-30B support it so far, with llama.cpp support via PR #27342.

Z-Lab's DFlash-2 boosts LLM throughput with diffusion-based drafting
#Z-Lab#Llama.cpp#Qwen#Muse-Glimmer
Read next
AI

Inception Labs launches Mercury 2.5, a diffusion LLM at 1,107 tokens/sec

AI

Cerebras claims 5x inference throughput gain via disaggregation

AI

QAT Gemma 4 26B-A4B on one TPU v6e: 15.6x KV cache, 1.9x throughput vs FP8

AI

Google Research Introduces Retrieve-for-Train (R4T): RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out