Open-source proxy with fine-tuned model cuts Codex context by 29.6%
A Show HN project offers a local proxy that intercepts Codex tool-call output and compresses it with a fine-tuned Qwen model, trimming tokens by 29.6% without invalidating the KV cache. The team was spending up to $700 per day per person on API calls due to bloated context.
- Proxy compresses tool output before it enters context, keeping KV cache valid
- 29.6% token reduction via fine-tuning on coding-agent trajectories
- Installs as a shell wrapper, disabled with codex --uncompress
- Risks: lost data, added latency, trust in the proxy binary
Read next
AI