chiprook
← Software
SoftwareOctober 9, 2026, 08:54

llama-server prompt cache reuses KV computed under a different LoRA scale

In llama.cpp b11514, llama-server's RAM prompt cache (--cache-ram, on by default) stores a slot's KV without recording the LoRA adapters and scales it was computed with. The LoRA-change check compares the new request against the slot's previous request rather than the restored entry, so answers are decoded on KV from another scale: 10 of 10 prompts were affected. Workarounds are --cache-ram 0 or cache_prompt: false.

llama-server prompt cache reuses KV computed under a different LoRA scale
#Llama.cpp
Read next
AI

Google TurboQuant compresses KV cache to 3 bits with no accuracy loss

Hardware

Huawei Debuts OceanStor M900: PB-Scale Shared KV Cache With ~60μs NPU-to-SSD Hop

Hardware

Huawei Connect 2026: OceanStor M900 turns PB-scale KV cache into AI infrastructure layer

Software

llama-server returns placeholder logprobs when speculative decoding is on