WeirdChat: A catalog of unexpected AI behaviors, discovered automatically

·LessWrong··

[This is a link-post for https://transluce.org/weirdchat. We recommend reading the website version for interactive visualizations.] Language models can behave in surprising and sometimes harmful ways. Yet as models have improved, these behaviors have become harder to find, often only appearing after widespread use. To surface these behaviors in simulation, we use automated techniques to elicit over 1,300 behavioral patterns in frontier open-weight models, some relatively benign, like making up a...

Read full article →

Related Articles

Apple defeats liability for not scanning iCloud for CSAM
speckx · Hacker News · 9h ago
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
piotrgrabowski · Hacker News · 1h ago
New US homeownership measure puts people first
throw0101a · Hacker News · 12h ago
FreeInk: Open ecosystem for e-readers
FriedPickles · Hacker News · 5h ago
Hacker wipes Romania's land registry database
speckx · Hacker News · 1d ago