GLM 5.3 API Cost: Always Thinking, 2.5x Cheaper per Answer Than 5.2
Z.ai released GLM 5.3 and GLM 5.3 Flash at the same price as GLM 5.2 ($1.40 per million input and $4.40 per million output tokens), but thinking mode cannot be disabled—only low, high, and max. On 33 verifiable tasks, GLM 5.3 in max mode scored 33/33 at $0.00468 per correct answer versus $0.01173 for GLM 5.2.
- GLM 5.3 Flash: 320B parameters, 18B active, $0.15/$0.50 per million tokens
- Disabling thinking returns error 1210
- Only low, high, and max (default) available
- GLM 5.3 in max: 33/33 correct, median 538 reasoning tokens vs 1129 for GLM 5.2
Read next
AI