BLOOM-WILT: on-policy examples of any LLM behaviour, from a one-line description and logits alone
Summary of my AI safety paper, BLOOM-WILT, written under the supervision of Edoardo Manino. Code is on GitHub, all transcripts are on HuggingFace, and LogitTilt is now available as part of Inspect.Reach out if you want to collaborate on a second paper building on this work. Thanks to BlueDot Impact for covering our compute expenses.TL;DRAutomated LLM auditors (like Anthropic's BLOOM) make red-teaming cheap to scale but their lack of optimisation pressure makes them inefficient at finding example...
Read full article →