WeirdChat: A catalog of unexpected AI behaviors, discovered automatically

·LessWrong··

[This is a link-post for https://transluce.org/weirdchat. We recommend reading the website version for interactive visualizations.] Language models can behave in surprising and sometimes harmful ways. Yet as models have improved, these behaviors have become harder to find, often only appearing after widespread use. To surface these behaviors in simulation, we use automated techniques to elicit over 1,300 behavioral patterns in frontier open-weight models, some relatively benign, like making up a...

Read full article →

Related Articles

Formalizing Fermat's Last Theorem
jlebar · Hacker News · 6h ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 3h ago
Can AI design circuit boards yet?
iopapa · Hacker News · 5h ago
Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
beardyw · Hacker News · 18h ago
Solving the Jane Street reverse engineering challenge
anitil · Hacker News · 15h ago