chiprook
← AI
AIOctober 1, 2026, 07:06

Open-source proxy with fine-tuned model cuts Codex context by 29.6%

A Show HN project offers a local proxy that intercepts Codex tool-call output and compresses it with a fine-tuned Qwen model, trimming tokens by 29.6% without invalidating the KV cache. The team was spending up to $700 per day per person on API calls due to bloated context.

Open-source proxy with fine-tuned model cuts Codex context by 29.6%
#OpenAI#Codex#Qwen
Read next
AI

Nokia open-sources AnyJev: a training-free layer that turns any open LLM into a calibrated decision model

AI

Open-source coding LLMs compared: GLM-5.3-Flash, Qwen3.8-Flash-Next, DeepSeek V4 Flash

AI

TaichuAI Open-Sources ZDTaichu5.0-9B Spatial Multimodal Model

AI

Xiaomi Open-Sources Robotics-U0 Embodied World Model and Training Stack