MLPerf Inference v6.1 adds end-to-end RAG and desktop agent benchmarks
MLCommons released MLPerf Inference v6.1 with two new tests: an End-to-End RAG benchmark for the full retrieval pipeline and an Edge Agentic benchmark replaying 1,007 turns of a coding agent on a desktop. A record 30 organizations submitted 120 systems and 486 results; DeepSeek-R1 is 5.7x faster than a year ago.
- 30 organizations submitted 120 systems and 486 results in MLPerf Inference v6.1
- DeepSeek-R1 is 5.7x faster year-over-year, VLMs 2.99x faster in six months
- RAG test uses 107,484 passages and 824 FRAMES tasks with a 97% accuracy target
- DGX Spark completed 1,007 agent turns in under 64 minutes at 20.1 tokens/s
Read next
Hardware