Modular Pretraining Enables Access Control

·LessWrong··

Full author list: Ethan Roland*, Murat Cubuktepe*, Erick Martinez*, Stijn Servaes, Keenan Pepper, Mike Vaiana, Diogo Schwerz de Lucena, Judd Rosenblatt, Addie Foote, Cem Anil, Alex Cloud; *Equal contributiontldr: Frontier AI models have knowledge that could be misused for nefarious purposes. To address this risk, we introduce Gradient Routed Auxiliary Modules (GRAM), a method for isolating dangerous knowledge to specific modules within a language model. These modules can be switched on or off to...

Read full article →

Related Articles

There's no reason for software to be slow anymore
Jach · Hacker News · 1d ago
hdiutil is deprecated in macOS 27 Golden Gate
zdw · Hacker News · 8h ago
Kobo can run apps now
thepoet · Hacker News · 1d ago
Malicious Rust crate Arrayref runs a build-time payload
abhisek · Hacker News · 2d ago
Z80 – The 1970s Microprocessor Still Alive (2021)
asdefghyk · Hacker News · 18h ago