Dust: Training AI Without Backpropagation
Researchers introduced Dust, a zeroth-order optimization method that pretrains neural networks without computing gradients. At large population sizes it matches backprop performance and is 10³–10⁴ times more efficient than weight-space evolutionary strategies.
- Dust replaces gradients with random perturbations of activations at each layer
- It is 10³–10⁴ times more efficient than weight-space evolutionary strategies
- Larger models proved more population-efficient than smaller ones
- Gradient estimates stay aligned with backprop up to 1B tokens
Read next
AI