chiprook
← AI
AIOctober 8, 2026, 23:32

Scality AI Inference Factory serves KV cache from object storage over RDMA

Scality launched AI Inference Factory, an open-code stack for running open-weight models on customer-owned infrastructure. The KV cache lives in ADI object storage and reaches GPUs over RDMA: context restore is 14x faster than recompute, with up to 8TB of cache and 1,000 resumable sessions.

Scality AI Inference Factory serves KV cache from object storage over RDMA
#Scality#VLLM#Nvidia
Read next
AI

Google TurboQuant compresses KV cache to 3 bits with no accuracy loss

AI

QAT Gemma 4 26B-A4B on one TPU v6e: 15.6x KV cache, 1.9x throughput vs FP8

AI

Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts

AI

vLLM 0.30.0 weight cache can serve another checkpoint's weights