TokenRouter: token-level LLM routing cuts inference cost 18% and latency 31%
TokenRouter is the first publicly available system that picks which language model generates each individual token instead of routing whole queries. Benchmarks show median per-token latency drops 31–32% and inference cost falls 18–21% with no quality loss.
- Median per-token latency falls from 3.4 ms to 2.3 ms on A100-40 GB
- GPU-seconds per 1M tokens drop from 1.02 s to 0.84 s (-21%)
- About 68% of tokens go to a cheap 7B model, only 32% hit the 70B model
- ROUGE-L for summarization rises 0.3%, BLEU for translation 0.1%
Read next
AI