SONI: Selective Orthogonalisation via Noise Injection

·LessWrong··

This project was completed as a capstone for TARA. All code is available in github.TL;DRThe Problem: Neural networks use superposition to pack many concepts into small latent spaces by making feature vectors almost-orthogonal. This entanglement makes models opaque and breaks safety interventions (e.g. concept erasure, activation steering) which rely on clean, isolated concept directions.The Gap: Full orthogonalisation (via sparsity penalties) destroys model capacity, while Sparse Autoencoders (S...

Read full article →

Related Articles

My security camera shipped a GitHub admin token in its login page
hhh · Hacker News · 16h ago
JEP 541: Deprecate the macOS/x64 Port for Removal
pmg1991 · Hacker News · 11h ago
DARPA, U.S. Air Force fly AI-controlled F-16
r2sk5t · Hacker News · 1d ago
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
adam_rida · Hacker News · 1d ago
Postgres LISTEN/NOTIFY actually scales
KraftyOne · Hacker News · 8h ago