BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens
BottleCap AI has released ThinkingCap-Qwen3.8-27B, a fine-tune of Qwen3.8-27B that cuts thinking tokens by an average of 37.2% across 12 benchmarks while macro-average accuracy drops just 0.86pp, from 86.65% to 85.79%. It is a drop-in replacement for Qwen3.8-27B on vLLM or SGLang, with FP8, NVFP4, GGUF and MLX builds.
- Mean thinking tokens fall from 15,735 to 12,144
- AA-LCR accuracy rises 2.25pp; AIME 2026 drops 3.85pp
- Builds: NVFP4 21 GB, GGUF 16-55 GB, MLX 4-bit 21 GB
- PolyForm Small Business license; commercial use needs an agreement
Read next
AI