"Correct Answer Features" Cannot Explain Multiple Choice Capabilities

·LessWrong··

TL;DR: I present theoretical and empirical evidence that LLMs cannot be (exclusively) using a "correct answer feature" as the main mechanism by which they perform multiple choice question answering. A hypothetical correct-answer feature would indicate the "correctness" of an option on the final token(s) of that option. However, such a mechanism cannot be used in all cases, and evidence from direct-effect head attribution indicates that a similar mechanism is used both in cases where a correct-an...

Read full article →

Related Articles

U.S. Strategic Petroleum Reserve Falls to Lowest Level Since 1982
thelastgallon · Hacker News · 7h ago
Does Reddit have an astroturfing problem? What the data suggests
p-s-v · Hacker News · 20h ago
Nvidia wants to put a watchdog chip next to every AI agent
jonbaer · Hacker News · 17h ago
Nissan's third generation e-POWER powertrain
mroche · Hacker News · 1d ago
MicroLLM Lab – Try 7 tiny LLM's in the browser
logicallee · Hacker News · 14h ago