chiprook
← AI
AISeptember 15, 2026, 21:13

Six Billion Requests Later: A Full Year of LLM Serving

A preprint, A Year in LLM Serving (arXiv:2608.13573), analyzes a year of production traffic on a serverless inference platform: 6.12 billion requests, 314,970 users, 9,174 models and 875,921 instances from April 2025 to April 2026. The authors plan to release the trace.

Six Billion Requests Later: A Full Year of LLM Serving
#DeepSeek#Qwen#MiniMax
Read next
AI

Google Gemini Notebook Brings Interactive Learning Overviews to All Users

AI

DOJ seeks AI legal assistant for pilot program

AI

Gemini Notebook adds configurable output languages

AI

Kev: open decision models let AI agents choose without generating text