New AI Training Method Matches Backpropagation Without Backprop

Dust: Pretraining Transformers Without Backpropagation

New AI Training Method Matches Backpropagation Without Backprop

Dust, a zeroth-order method, perturbs activations independently at every token, turning each token into a virtual population member. This lets one forward pass evaluate thousands of candidates in parallel, making it up to 10,000 times more efficient than weight-space evolution strategies like EGGROLL. Surprisingly, larger models are more population-efficient, and at scale Dust can exceed backpropagation, hinting that brute-force search may rival gradient-based learning in compute-rich regimes.

Dust approximates backprop closely at large population (i.e. substantially more compute) and in multiple settings even exceeds it. This hints that in a compute-rich regime we might be able to surpass backprop.

More from this day

2026-10-05