llama-server prompt cache reuses KV computed under a different LoRA scale
In llama.cpp b11514, llama-server's RAM prompt cache (--cache-ram, on by default) stores a slot's KV without recording the LoRA adapters and scales it was computed with. The LoRA-change check compares the new request against the slot's previous request rather than the restored entry, so answers are decoded on KV from another scale: 10 of 10 prompts were affected. Workarounds are --cache-ram 0 or cache_prompt: false.
- Cache entries store only tokens and checkpoints, no LoRA or scale info
- On b11514 the bug reproduced on 10 of 10 prompts; b10703 behaved the same
- With --cache-ram 0 or cache_prompt: false, 10 of 10 answers were correct
- With -np 4 each prompt keeps its own slot and the bug does not appear
Read next
AI