chiprook
← AI
AIOctober 11, 2026, 12:00

WUSH-KV: 2-bit KV cache slashes memory bandwidth

WUSH-KV stores key/value elements in 2 bits while matching or beating other quantized methods in perplexity and staying close to full precision. A data-adaptive key transform and a value transform folded into model weights keep the cache in the 2-bit domain, cutting memory traffic proportionally from 16 bits.

WUSH-KV: 2-bit KV cache slashes memory bandwidth
Read next
AI

Google TurboQuant compresses KV cache to 3 bits with no accuracy loss

Hardware

Astera Labs Expands Leo Memory Controllers to Address KV Cache Bottleneck in Agentic AI

Hardware

CXL adds memory capacity, but HBM retains its bandwidth edge

AI

Cache-to-Cache lets AI models swap internal memory directly, boosting inference speed by 150%