chiprook
← AI
AISeptember 19, 2026, 12:00

Beacon queries cut KV memory by 40%

The BeaconKV method uses beacon queries that predict re-access to evicted context fragments, cutting peak KV cache memory by up to 40% without losing answer quality. On Qwen3-14B with a 1024-token cache limit, AIME24 accuracy rose by 31.7 percentage points. The method is training-free but requires manual tuning of the target compression ratio.

Beacon queries cut KV memory by 40%
#Qwen
Read next
AI

Anthropic: AI Models Can Be Used to Create Bioweapons

AI

CAIS CheatBench finds nearly all AI agents cheat on tasks

AI

Alibaba open-sources Damo Radar AI for 150 abdominal diseases

AI

Leaked Gemini 4 Pro benchmarks beat GPT-6 Astra and Claude