Refusal Is Complicated As Hell: An Update

·LessWrong··

TL;DRIt would make sense to briefly skim through our previous post that introduces our experiments on refusal in LLMs. There we explain how it started, here we’ll tell how it’s going.The primary goal of this text is to try and structure the list of whack-a-mole research questions. The secondary goal is to get some outside perspective, so if you run a similar research or have seen a similar research, please lend us a hand.Feel free to jump straight to the section that looks most appealing. We rec...

Read full article →

Related Articles

Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 20h ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 1d ago
Ask HN: Who's still keeping a DOS machine up because the business depends on it?
mlaux · Hacker News · 21h ago
Understanding the Impact of LLM Watermarking on AI Agent Behavior
nisosguy · Hacker News · 4h ago
Excel now supports multiple values in a single cell
luispa · Hacker News · 20h ago