Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible

·LessWrong··

Authors: Kyle O’Brien¹, Spencer Kitts², Cameron Tice¹, Alek Westover² ¹Geodesic Research, ²Redwood ResearchFigure 1: We pretrained LLMs from scratch with and without filtering subversion-relevant information, finding that subversion-relevant knowledge is significantly diminished while retaining general ML knowledge.TLDR: Filtering subversion-relevant information — subversion strategies, information about subversion defences, and empirical evaluations of the effectiveness of subversion strategies...

Read full article →

Related Articles

Pixel 11 doesn't yet meet the GrapheneOS security standards and may be skipped
finnlab · Hacker News · 6h ago
Improper redaction reveals Google Data Center water and electricity usage
sensanaty · Hacker News · 1d ago
US closely monitoring case of lab worker who possibly died of plague in Siberia
tosh · Hacker News · 3h ago
Mold Linker Version 3.0.0 Release – Rewritten in Rust
roflcopter69 · Hacker News · 8h ago
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
snehesht · Hacker News · 1d ago