chiprook
← AI
AIOctober 3, 2026, 14:45

vLLM 0.30.0 weight cache can serve another checkpoint's weights

In vLLM 0.30.0 the weight cache (load_format="ipc_cache") fingerprints checkpoints only by shard file names and safetensors headers, not tensor values. A base model and its fine-tune with identical layout get the same key, so an engine can silently run on another checkpoint's weights.

vLLM 0.30.0 weight cache can serve another checkpoint's weights
#VLLM#Qwen
Read next
AI

Repacked QAT Gemma 4 on one TPU v5e: 12B serves at 675 tokens/s

AI

QAT Gemma 4 26B-A4B on one TPU v6e: 15.6x KV cache, 1.9x throughput vs FP8

AI

Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts

AI

Six Billion Requests Later: A Full Year of LLM Serving