Do LoRA Read Directions Encode Visual Concepts?

·LessWrong··

TLDRI compare the semantic coherence of read directions learned by standard, ReLU, and TopK LoRA adapters with random directions in CLIP’s residual stream. Clarity, a measure of semantic coherence, is concentrated at the positive and negative extremes of the activation distribution.Random directions can occasionally produce highly coherent examples, so a convincing activation grid alone does not show that a concept was learned. However, learned directions are more consistently coherent: 86% of s...

Read full article →

Related Articles

Document-borne AI worms can self-propagate through Copilot for Word
Canopy9560 · Hacker News · 3h ago
Handbook.md shows that long policy documents do not reliably govern agents
spIrr · Hacker News · 1h ago
New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 1d ago
US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 2d ago
Zig's Incremental Compilation Internals
garyhtou · Hacker News · 23h ago