chiprook
← AI
AISeptember 19, 2026, 22:12

DeepSeek Cuts AI Agent Memory Cost 4x: New Architecture Fits More Sessions Per GPU

DeepSeek released the V4.1-Flash model with a Causal Encoder-Decoder architecture that shares cached states across transformer layers. GPU memory usage per AI agent session is reduced by 75%, allowing four times as many concurrent sessions on a single accelerator. Weights are available under the MIT license, but are subject to Chinese intelligence law.

DeepSeek Cuts AI Agent Memory Cost 4x: New Architecture Fits More Sessions Per GPU
#DeepSeek
Read next
AI

Anthropic: AI Models Can Be Used to Create Bioweapons

AI

CAIS CheatBench finds nearly all AI agents cheat on tasks

AI

Alibaba open-sources Damo Radar AI for 150 abdominal diseases

AI

Leaked Gemini 4 Pro benchmarks beat GPT-6 Astra and Claude