LLM post-training digest: GRPO moves to OCR, DPO and GRPO get debugged
A digest of LLM post-training research from October 2–9, 2026 covers PEFT, preference optimization (DPO/GRPO/RLHF), distillation and synthetic data. Highlights include LightOnOCR-3 applying GRPO to OCR (86.3 on olmOCR-Bench), DIAL-OPD beating full-token training with just 40% of tokens, and a 270M Falcon OCR Arabic model ranking #2 of 17.
- LightOnOCR-3 applies GRPO to OCR: 86.3 on olmOCR-Bench, leads 4B/0.8B tiers
- DIAL-OPD: training on 40% of tokens yields up to +5.25pp on math reasoning
- Falcon OCR Arabic at 270M params ranks #2 of 17 via SFT+RL
- SGUID selects a skill bank up to 11x smaller with no quality loss
Read next
AI