"Correct Answer Features" Cannot Explain Multiple Choice Capabilities

·LessWrong··

TL;DR: I present theoretical and empirical evidence that LLMs cannot be (exclusively) using a "correct answer feature" as the main mechanism by which they perform multiple choice question answering. A hypothetical correct-answer feature would indicate the "correctness" of an option on the final token(s) of that option. However, such a mechanism cannot be used in all cases, and evidence from direct-effect head attribution indicates that a similar mechanism is used both in cases where a correct-an...

Read full article →

Related Articles

GLM-5.3: Frontier coding with emergent cyber capabilities
pella · Hacker News · 21h ago
Firefox is now the last major browser that still supports uBlock Origin
DemiGuru · Hacker News · 8h ago
Going Dark, and the era of law enforcement hacking
vslira · Hacker News · 6h ago
In Australia, a home battery boom has helped cut wholesale power prices
speckx · Hacker News · 13h ago
Where did the old web go? We followed 657,607 links to find out
tdx · Hacker News · 1d ago