Counterfactual Resampling to Analyse Model Behaviour

·LessWrong··

Done as my final project for a BlueDot Impact Technical AI Safety sprint, facilitated by BAISH (Buenos Aires AI Safety Hub)TL;DR: I measured Value Leakage (Betley et al.) per conversation, cutting CoTs at different points and resampling under a prompt that flips which answer serves the model's values.Betley et al.’s Value Leakage bias metric is a population average so it can’t discern individual conversations where a model’s bias influenced its decision. I designed a method to measure the same b...

Read full article →

Related Articles

Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 3h ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 3h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
Formalizing Fermat's Last Theorem
jlebar · Hacker News · 1d ago
Why are European countries moving their gold out of North America?
ranit · Hacker News · 17h ago