Six Billion Requests Later: A Full Year of LLM Serving
A preprint, A Year in LLM Serving (arXiv:2608.13573), analyzes a year of production traffic on a serverless inference platform: 6.12 billion requests, 314,970 users, 9,174 models and 875,921 instances from April 2025 to April 2026. The authors plan to release the trace.
- 6.12 billion requests, 35.8T input and 2.5T output tokens in a year
- Daily active models rose from under 100 to over 400 at peak
- Median response length fell from several hundred tokens to under 100
- FIFO and LRU match complex cache algorithms due to strong temporal locality
Read next
AI