Headroom: context compression cuts agent token costs by 60–95%
Headroom is a compression layer between agents and LLMs: the kompress-v2-base model shrinks tool outputs, logs, RAG chunks and files. It claims 20% token savings for coding agents and 60–95% for JSON-heavy workflows without changing answers. It ships as a Python library, FastAPI proxy and MCP server.
- 20% token savings for coding agents, 60–95% for JSON-heavy workflows
- A 55,957-token prompt compressed to 24,340 with the FATAL line preserved
- Three modes: library (50–200ms), proxy (100–300ms), MCP (150–400ms)
- Model preserves function signatures, stack traces and ERROR/FATAL lines
Read next
Hardware