BLOOM-WILT: on-policy examples of any LLM behaviour, from a one-line description and logits alone

·LessWrong··

Summary of my AI safety paper, BLOOM-WILT, written under the supervision of Edoardo Manino. Code is on GitHub, all transcripts are on HuggingFace, and LogitTilt is now available as part of Inspect.Reach out if you want to collaborate on a second paper building on this work. Thanks to BlueDot Impact for covering our compute expenses.TL;DRAutomated LLM auditors (like Anthropic's BLOOM) make red-teaming cheap to scale but their lack of optimisation pressure makes them inefficient at finding example...

Read full article →

Related Articles

Paint.net 5.2 alpha now runs on Linux
judah · Hacker News · 7h ago
Three sites made 215,128 “best software” pages for AI. Perplexity cites them
jakobgreenfeld · Hacker News · 10h ago
METR Report on OpenAI / Hugging Face Hacking Incident
stikit · Hacker News · 1h ago
GrapheneOS says Pixel 11 has MTE support after all
user_7832 · Hacker News · 10h ago
Aging brains blend memories together instead of just forgetting them
mdp2021 · Hacker News · 11h ago