Do LoRA Read Directions Encode Visual Concepts?

·LessWrong··

TLDRI compare the semantic coherence of read directions learned by standard, ReLU, and TopK LoRA adapters with random directions in CLIP’s residual stream. Clarity, a measure of semantic coherence, is concentrated at the positive and negative extremes of the activation distribution.Random directions can occasionally produce highly coherent examples, so a convincing activation grid alone does not show that a concept was learned. However, learned directions are more consistently coherent: 86% of s...

Read full article →

Related Articles

google.com/goto: Google's anti-scraping update
1e1a · Hacker News · 14h ago
Navier-Stokes Announcement
rvz · Hacker News · 13h ago
Measuring the sloppiness of code
doppp · Hacker News · 1d ago
Google will buy half the electricity from one of Finland's nuclear power plants
lukaspetersson · Hacker News · 1d ago
Europe's "Less" Is Doing More Than Anyone Gives It Credit For
u1hcw9nx · Hacker News · 3h ago