Dust: Pretraining Transformers Without Backpropagation

·Hacker News··

Contents TL;DR 1 Introduction 2 Method 2.1 Activation-Space Perturbation 2.2 Credit Assignment 2.3 Interference and Tuning 3 Pretraining Without a Backward Pass 3.1 Setup 3.2 Main Results 3.3 Dust Under Adam 4 Search in High-Dimensional Space 4.1 Overparameterization 4.2 Emergence of Backprop-Like Gradients 5 Conclusion 6 Related Work References Appendix Q Labs Research Dust: Pretraining Transformers Without Backpropagation Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna October 202

Read full article →

Related Articles

Opus 5.5 agents discover two room-temperature magnetic semiconductor candidates
outlier99 · Hacker News · 5h ago
Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped
finnlab · Hacker News · 13h ago
US closely monitoring case of lab worker who possibly died of plague in Siberia
tosh · Hacker News · 9h ago
Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 1d ago
Mold Linker Version 3.0.0 Release – Rewritten in Rust
roflcopter69 · Hacker News · 14h ago