When talked into harm, a model blames the answer, not itself (an interpretability study of guilt vs shame)

·LessWrong··

This was my write-up for Neel Nanda’s Winter 2027 MATS Stream (~20h research task). I didn’t get in, but it was my first application, so there’s always next time :P Lightly restructured here to fit the LessWrong format better. Repo: https://github.com/star2vec/guiltea── ⋆⋅☆⋅⋆ ──TL;DR: Models are safety-trained, but they can still be persuaded to do harmful acts. When that happens, how does blaming or informing it of its mistake influence its understanding of itself, its role, and subsequent acti...

Read full article →

Related Articles

NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 11h ago
What happened to the Snowden archive
EXHades · Hacker News · 1d ago
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 1d ago
Ask HN: Is it impossible to disable Siri on macOS 27?
semidror · Hacker News · 17h ago
HERMES radio enables voice and data communication over vast distances
SamuraiLion · Hacker News · 14h ago