WeirdChat: A catalog of unexpected AI behaviors, discovered automatically
[This is a link-post for https://transluce.org/weirdchat. We recommend reading the website version for interactive visualizations.] Language models can behave in surprising and sometimes harmful ways. Yet as models have improved, these behaviors have become harder to find, often only appearing after widespread use. To surface these behaviors in simulation, we use automated techniques to elicit over 1,300 behavioral patterns in frontier open-weight models, some relatively benign, like making up a...
Read full article →