Predicting LLM Safety Before Release by Simulating Deployment

·LessWrong··

Paper linkBefore releasing a new model, labs need to understand not just what it can do, but how it is likely to behave in real-world use, including where it might introduce new risks. This becomes even more important as capabilities increase. As part of our pre-deployment safety review, we leverage targeted evaluations, red-teaming, and other checks to understand model behavior. We’ve now started using a method for simulating model deployments before they happen, which adds a complementary sign...

Read full article →

Related Articles

Apple is about to make Hide My Email useless
SXX · Hacker News · 14h ago
TIL: You can make HTTP requests without curl using Bash /dev/TCP
mrshu · Hacker News · 16h ago
A backdoor in a LinkedIn job offer
lwhsiao · Hacker News · 1d ago
Mechanical Watch (2022)
razin · Hacker News · 21h ago
Google Chrome's Next Update Will Mark the End of Popular Ad Blockers
arnejenssen · Hacker News · 17h ago