OrcaSAQ-2 Shrinks 27B Qwen 3.8 AI Into a 12.3GB Local Model
OrcaSAQ-2 is a quantized version of the 27-billion-parameter Qwen 3.8, compressed from 54GB to 12.3GB while retaining 93.2% token agreement on WikiText-2. It needs a GPU with at least 16GB of memory, solves 70% of S.S.S. swbench tasks and scores 58.4 on Terminal Bench 2.1.
- Compressed from 54GB to 12.3GB via quantization with 93.2% token agreement
- Requires a GPU with at least 16GB memory, 32,000-token context recommended
- Solved 70% of S.S.S. swbench tasks and scored 58.4 on Terminal Bench 2.1
- Speculative decoding boosts throughput but strains memory
Read next
AI