VFUSE: Virulent Feature Understanding With Sparse AutoEncoders

·LessWrong··

AbstractGenerative models have shown remarkable progress in a variety of domains such as protein design, but such power enables the opaque generation of hazardous proteins. In this work, we introduce VFUSE (Virulent Feature Understanding with Sparse autoEncoders), a mechanistic interpretability approach that trains SAEs on diffusion-transformer activations to audit protein models for hazard-aware features. We apply VFUSE to RoseTTAFold3 and RFDiffusion3, popular open-weight models for protein fo...

Read full article →

Related Articles

GCC steering committee announces AI policy
arto · Hacker News · 2h ago
AI's top startups are barely publishing their research
YeGoblynQueenne · Hacker News · 16h ago
Document-borne AI worms can self-propagate through Copilot for Word
Canopy9560 · Hacker News · 1d ago
Keychron announces first open-source firmware for gaming mice
JLO64 · Hacker News · 21h ago
Turning a dumb AC unit smart (without losing my security deposit)
austinallegro · Hacker News · 19h ago