Meta releases Llama 4 Scout: 105B parameters with 1M-token context
Meta FAIR released Llama 4 Scout, an open-weights model built on Mixture-of-Depths routing with 105B total and 24B active parameters per token. It natively supports a 1M-token context with 99.8% NIAH retrieval accuracy and runs in 4-bit mode on 24GB VRAM. Checkpoints are on Hugging Face, with support in vLLM, Ollama and LM Studio.
- 105B total parameters, 24B active per token
- Native 1M-token context with 99.8% NIAH accuracy
- Runs in 4-bit mode on 24GB VRAM (RTX 4090/5090)
- Resolves 48.6% of tasks on SWE-bench Verified
Read next
AI