Banana in, Bostrom out: paperclip maximization is one token-direction swap away (in Qwen 3.6-27B)
IntroductionI started learning about interpretability late February of this year. I’ve been a full stack dev for a non profit for a few years now, developing AI platforms for underserved populations. But I had never taken a look at the inside of the machinery of the systems I was using to build the platforms.After Anthropic released the Global Workspace paper a couple weeks ago, I started poking around the Neuronpedia UI J Space lens for Qwen3.6-27B before switching to API calls. What I found wa...
Read full article →