chiprook
← AI
AIOctober 2, 2026, 17:13

ISOM-R2 streams 1,055,402 tokens at 3.24 GB peak VRAM

The new recurrent memory architecture ISOM-R2 (Isometric State Space / Virtual SVD) ingested a 1,055,402-token context of real Python code on the Aazhi-Coder-1.5B model while holding peak VRAM at 3.24 GB — 8.7x less than the 28.2 GiB a standard KV cache would need. Throughput reached about 8,526 tokens/sec with 100% macro retrieval precision.

ISOM-R2 streams 1,055,402 tokens at 3.24 GB peak VRAM
#ISOM-R2#Qwen2.5-Coder#HuggingFace
Read next
Hardware

Nvidia updates RTX Mega Geometry to 2.0 with on-demand geometry streaming to VRAM

AI

Perplexity’s local AI agent comes to Windows, but only for RTX GPUs with at least 24GB of VRAM

AI

OrcaSAQ-2 Shrinks 27B Qwen 3.8 AI Into a 12.3GB Local Model

AI

OpenAI's Dots VM Specs Leak: AMD EPYC 9V74, 9.73 GB RAM Per Instance