chiprook
← AI
AISeptember 22, 2026, 10:44

Inception Labs' Mercury 2.5 named fastest LLM at 1,107 tokens/sec

Inception Labs says its diffusion model Mercury 2.5 hits 1,107 tokens per second on Nvidia GPUs at $0.20/$0.75 per million tokens. OpenRouter measures 440 tok/s at P50 — still faster than Gemini 3.5 Flash-Lite, GPT-5.6 Luna and Claude Haiku 4.5. The 80% launch discount ended on September 8, 2026.

Inception Labs' Mercury 2.5 named fastest LLM at 1,107 tokens/sec
#InceptionLabs#Mercury#OpenAI#Anthropic
Read next
AI

Kyutai releases Voice of Reason: speech-native models solve spoken math

AI

Instinct founder Noah Shinn reportedly raising at $10B valuation

AI

Ex-Google safety chief warns AI could harm children more than social media

AI

Experts debate risk of an AI takeover of the internet