ISOM-R2 streams 1,055,402 tokens at 3.24 GB peak VRAM
The new recurrent memory architecture ISOM-R2 (Isometric State Space / Virtual SVD) ingested a 1,055,402-token context of real Python code on the Aazhi-Coder-1.5B model while holding peak VRAM at 3.24 GB — 8.7x less than the 28.2 GiB a standard KV cache would need. Throughput reached about 8,526 tokens/sec with 100% macro retrieval precision.
- 1,055,402-token context from 181 transformers repo files processed in 123.78 seconds
- Peak VRAM 3.24 GB vs 28.2 GiB for a standard KV cache — 8.7x lower
- Prefill stayed flat at 2.97 GB across all 1M tokens, active buffer under 400 MB
- 100% retrieval precision: only 2 target files isolated from 516 chunks
Read next
Hardware