Researchers say compress tool output, not agent actions, in coding agents
A Peking University paper (arXiv:2609.31430) shows that compressing an entire agent transcript breaks coding agents by blurring exact file paths and tool syntax. The LOHA layout keeps agent actions and the latest K observations in raw text while compressing older tool output into latent tokens, with Anchored Context Distillation preserving quality on plain text.
- 43% fewer tokens on Qwen3-4B and 57% on SWE-Master-4B-RL at K=3
- Resolve rate 12.1% vs 14.5% uncompressed for Qwen3-4B
- Under a 32K budget Qwen3 with K=3 resolved 21.1% vs 11.1%
- Single-GPU instance throughput improved 1.9x
Read next
AI