Inception Labs launches Mercury 2.5, a diffusion LLM at 1,107 tokens/sec
Inception Labs announced Mercury 2.5, billed as the largest diffusion language model ever trained, running at 1,107 tokens per second on widely available Nvidia GPUs. It claims a 40% intelligence gain over Mercury 2, matches cost-optimized frontier tier quality, and ships via an OpenAI-compatible API with a 260K-token context window.
- 1,107 tokens/sec on Nvidia GPUs vs 277.5 for Gemini 3.8 Flash
- Quality positioned against GPT-5.6 Luna Low, Gemini 3.5 Flash-Lite and Claude Haiku 4.5
- 260K-token context; $0.20/M input and $0.75/M output tokens
- Mercury Voice (sub-170ms TTFT) and Mercury Router also announced
Read next
AI