Research Note: Filtering Subversion-Relevant Information From Pretraining Data Is Feasible
Authors: Kyle O’Brien¹, Spencer Kitts², Cameron Tice¹, Alek Westover² ¹Geodesic Research, ²Redwood ResearchFigure 1: We pretrained LLMs from scratch with and without filtering subversion-relevant information, finding that subversion-relevant knowledge is significantly diminished while retaining general ML knowledge.TLDR: Filtering subversion-relevant information — subversion strategies, information about subversion defences, and empirical evaluations of the effectiveness of subversion strategies...
Read full article →