Modular Pretraining Enables Access Control

·LessWrong··

Full author list: Ethan Roland*, Murat Cubuktepe*, Erick Martinez*, Stijn Servaes, Keenan Pepper, Mike Vaiana, Diogo Schwerz de Lucena, Judd Rosenblatt, Addie Foote, Cem Anil, Alex Cloud; *Equal contributiontldr: Frontier AI models have knowledge that could be misused for nefarious purposes. To address this risk, we introduce Gradient Routed Auxiliary Modules (GRAM), a method for isolating dangerous knowledge to specific modules within a language model. These modules can be switched on or off to...

Read full article →

Related Articles

Mistral Large 4
Philpax · Hacker News · 18h ago
JetBrains reported a net financial loss first time in its tracked history
thw_9a83c · Hacker News · 20h ago
OpenTPU – An open-source AI accelerator, developed by AI
fsbonetto · Hacker News · 15h ago
AnyPS5: Port PS5 binaries to PC without emulation (87% system libraries mapped)
Fe2O3 · Hacker News · 8h ago
Nobel Prize in Physics 2026: Francis Halzen
solarist · Hacker News · 22h ago