Get the latest tech news
Dust: Pretraining Transformers Without Backpropagation
@misc{dahal2026backprop, title = {Dust: Pretraining Transformers Without Backpropagation}, author = {Dahal, Samip and Mandal, Bishwas and G{\"u}lbahar, Serdar and Vegesna, Akshay}, year = {2026}, url = {https://qlabs.sh/research/dust} } TL;DR - We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models. Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel.
None
Or read this on Hacker News

