Reading into VLM hallucinations using the Jacobian lens

·LessWrong··

Reading and editing the visual workspace of a vision-language model.Vision-language models hallucinate: ask one whether some object is in a picture and it will happily say yes whether it's there or not. Why is this?Using Anthropic's recent J-lens method I found that LLaVA-1.5-7B's internal state seems to register that the object is absent. The exact same evidence that the yes/no question ignores produces almost perfect answers when posed as a choice instead ("is this a lamp or a dog?").Perhaps t...

Read full article →

Related Articles

Coconut oil jet fuel matches kerosene's efficiency in engine tests
mdp2021 · Hacker News · 16h ago
Slovakia finds Russian backdoor in traffic speed cameras
dredmorbius · Hacker News · 17h ago
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
ed-is-ai · Hacker News · 15h ago
There's no reason for software to be slow anymore
Jach · Hacker News · 2d ago
How Complex Systems Fail (1998)
shortcrct · Hacker News · 16h ago