Activation Oracles significantly underperform without a safe base model
Raffaello Fornasiere*, Nikita Menon, Andrzej Szablewski, Gabriel Konar-Steenberg, Stefan Heimersheim.*First Author. This study is a focused extension of a project done at LASR Labs.Thanks to (in alphabetical order) Adam Karvonen, Alejandro Wainstock, Damiano Fornasiere, and Daniele Pace for discussions, thoughts, and reviews.TL;DRIn this study, we show that when an Activation Oracle (AO) is trained on a base model that already presents some undesirable behaviour, the AO becomes unreliable to ide...
Read full article →