Nvidia's LoGRA cuts LLM RL training memory by up to 45.7%
Nvidia and collaborators introduced LoGRA, which replaces full gradient buffers with low-rank sketches and adds predicted-KL step control. Average training memory drops by up to 45.7%, and a 27B model trains stably for over 1,100 steps on a single eight-GPU node.
- 7B model memory falls from 31.82 GiB to 17.29 GiB, a 45.7% reduction
- 27B model trains 1,100+ steps on one 8-GPU node where dense Adam runs out of memory
- 27B held-out macro score improves from 39.69% to 62.94%
- For 1.5B, LoGRA uses 7.18 GiB versus LoRA's 13.21 GiB
Read next
AI