Activation Oracles significantly underperform without a safe base model

·LessWrong··

Raffaello Fornasiere*, Nikita Menon, Andrzej Szablewski, Gabriel Konar-Steenberg, Stefan Heimersheim.*First Author. This study is a focused extension of a project done at LASR Labs.Thanks to (in alphabetical order) Adam Karvonen, Alejandro Wainstock, Damiano Fornasiere, and Daniele Pace for discussions, thoughts, and reviews.TL;DRIn this study, we show that when an Activation Oracle (AO) is trained on a base model that already presents some undesirable behaviour, the AO becomes unreliable to ide...

Read full article →

Related Articles

Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
TangerineDream · Hacker News · 9h ago
We found a division by zero bug in FFmpeg with a vibecoded fuzzer
dclavijo · Hacker News · 9h ago
Tell HN: PayPal Blocks GrapheneOS
leumon · Hacker News · 17h ago
Autism mutations drive neurodevelopmental pathology
slantedview · Hacker News · 8h ago
Decompiling a Nintendo 64 game in 84 days
knackers · Hacker News · 12h ago