Byte-level models beat tokenized ones in the long run, study finds
A University of Washington and Meta FAIR paper shows tokenized models sprint early then plateau, while byte-level models keep climbing. Distilling a 1B byte-level student from Llama 3-8B via End-of-Token conversion yields 4% gains at the compute asymptote with 1/6 the data and 1/5 the teacher-logit storage.
- Byte-distilled models lead token-distilled ones by 4% at the asymptote
- End-of-Token beats Marginalize-It by 1.9% via bijective mapping
- Byte student matches token model on 1/6 of the training data
- A 257-symbol byte vocabulary cuts teacher-logit storage to 1/5
Read next
AI