laya-compactor cuts RAG context tokens by 70% with no answer-quality loss
An open-source tool, laya-compactor, scores and filters retrieved documents locally in a single forward pass before the LLM runs. On SQuAD it kept exact match identical (0.345) while cutting tokens by 69.7%; on HotpotQA it retained 94.5% of gold documents with 70.3% token savings.
- 70% fewer input tokens with identical SQuAD exact match (0.345)
- 94.5% of gold documents retained on HotpotQA, EM dropped from 0.230 to 0.200
- Runs locally for free, unlike OpenAI Decisions API and TypeSafe Jev
- Drop-in integrations for LangChain and LlamaIndex, install via pip install laya-compactor
Read next
AI