Speeding Up JumpReLU SAE Inference with Custom Triton Kernels (2–14× on Real SAEs)

·LessWrong··

MotivationSparse Autoencoders (SAEs) have become a central tool in mechanistic interpretability research, providing a way to decompose a model's internal activations into sparse, interpretable features. However, extracting these features often requires running the SAE over large volumes of activations across many layers and tokens. This makes SAE inference efficiency a practical bottleneck for interpretability research at scale. This post focuses on improving the inference efficiency of JumpReLU...

Read full article →

Related Articles

google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 7h ago
Navier-Stokes Announcement
rvz · Hacker News · 6h ago
Measuring the sloppiness of code
doppp · Hacker News · 21h ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
The Deathray: A simple way for an untrusted site to freeze a Mac
auberonedu · Hacker News · 1d ago