Vast uses tiered storage to ease AI agent memory demands
Vast outlined a tiered memory approach for AI agents: KV cache is offloaded from GPU memory to CPU memory and persistent media, allowing petabytes of context to be retained and avoiding repeated recalculation. Nvidia's Dynamo software orchestrates the process, and a confidential computing service for sensitive workloads has also launched.
- A 500,000-token session consumes 1/10 to 1/20 of a GPU's memory
- KV cache is offloaded to CPU memory and persistent media up to petabytes
- Nvidia Dynamo software orchestrates the offloading process
- A confidential computing service for sensitive workloads has launched
Read next
AI