chiprook
← Software
SoftwareOctober 6, 2026, 10:37

llama-server returns placeholder logprobs when speculative decoding is on

With speculative decoding enabled (draft model, MTP or n-gram types), llama-server returns logprob 0.0 and an empty top_logprobs list for every token after the first. Mean logprob of 64-token samples at temperature 1 drops from -0.48 to -0.0011, while the generated text looks normal. The bug reproduces on release b11430 and earlier builds.

llama-server returns placeholder logprobs when speculative decoding is on
#Llama.cpp
Read next
Software

llama-server sleep race hangs or crashes requests arriving just before sleep

AI

AgSpec fixes speculative decoding for coding agents

AI

Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs

Security

ZoomEye: "Zimbra Web Client" body search returns 537,391 assets, almost none are mail servers