Banana in, Bostrom out: paperclip maximization is one token-direction swap away (in Qwen 3.6-27B)

·LessWrong··

IntroductionI started learning about interpretability late February of this year. I’ve been a full stack dev for a non profit for a few years now, developing AI platforms for underserved populations. But I had never taken a look at the inside of the machinery of the systems I was using to build the platforms.After Anthropic released the Global Workspace paper a couple weeks ago, I started poking around the Neuronpedia UI J Space lens for Qwen3.6-27B before switching to API calls. What I found wa...

Read full article →

Related Articles

Hackers Had a Live Feed of Every ID Verification Company Scanned for over a Year
beardyw · Hacker News · 7h ago
Solving the Jane Street Reverse Engineering Challenge
anitil · Hacker News · 4h ago
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
screm · Hacker News · 17h ago
OpenAI's GPT-6 Astra on ARC-AGI-3
vignesh_warar · Hacker News · 18h ago
Carbon-aware electricity pricing, measured daily on 38 grids
High-Five · Hacker News · 6h ago