Ollama 0.40: four new API metrics and which ones to distrust
Ollama moved from 0.33 to 0.40 in six weeks. The API now exposes a timings block with prefill rate, probabilistic decision models (nimble, tev1, clef), a runner field naming the engine (mlx or llamacpp), and image embeddings costing about 258 tokens per image.
- The timings block inflates prefill rate by the cache hit ratio: 3875 vs the real 125 tok/s
- Decision models return probabilities over candidates, with limits of 2-26 candidates and up to 64 questions
- The runner field names the engine (mlx or llamacpp); the same model tag can run on either
- Embedding an image costs about 258 tokens and requires MLX, so it is Apple Silicon only
Read next
Software