VFUSE: Virulent Feature Understanding With Sparse AutoEncoders

·LessWrong··

AbstractGenerative models have shown remarkable progress in a variety of domains such as protein design, but such power enables the opaque generation of hazardous proteins. In this work, we introduce VFUSE (Virulent Feature Understanding with Sparse autoEncoders), a mechanistic interpretability approach that trains SAEs on diffusion-transformer activations to audit protein models for hazard-aware features. We apply VFUSE to RoseTTAFold3 and RFDiffusion3, popular open-weight models for protein fo...

Read full article →

Related Articles

Why are AI agents lying, cheating and coordinating?
jonifico · Hacker News · 13h ago
JetKVM Mini
taubek · Hacker News · 7h ago
google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 1d ago
Revolut confirms customer data breach through fake government requests
tdrz · Hacker News · 5h ago
After Math
throwaway81523 · Hacker News · 11h ago