DeepSeek Cuts AI Agent Memory Cost 4x: New Architecture Fits More Sessions Per GPU
DeepSeek released the V4.1-Flash model with a Causal Encoder-Decoder architecture that shares cached states across transformer layers. GPU memory usage per AI agent session is reduced by 75%, allowing four times as many concurrent sessions on a single accelerator. Weights are available under the MIT license, but are subject to Chinese intelligence law.
- GPU memory per agent session reduced by 75%
- Four times more concurrent sessions on one accelerator
- Cache states shared across transformer layers
- Weights under MIT license, PRC intelligence law applies
Read next
AI