The J-lens offset is the model's token frequency: z-scoring helps
This is a linkpost for the write-up on my site; the full body is below, and the code, decisions ledger and devlog are in the repo. Base-model z-score calibration of the J-lens helps elicit hidden secret words from Cywiński et al.'s taboo organisms: 0.805 leave-one-out accuracy against 0.665 for their protocol on Gemma-2-9B-it, and the only non-zero readout on Qwen3-1.7B. The J-vs-logit part of that gap is a point estimate at n = 20 (paired sign-flip p ≈ 0.19; p ≈ 0.23 against a z-scored logit le...
Read full article →