Reasoning and learning about injected concepts in language models

·LessWrong··

This work was done as a part of SPAR, under the mentorship of Mirko Bronzi and Damiano Fornasiere. TL;DRWe test models' ability to recover information about their activations by injecting steering vectors, and asking the LLMs to verbalize properties of them. We train models with in-context learning and test for three capabilities: Can models identify the region of layers (early, middle, late) of the injection?Can models identify the relative magnitude (low, medium, high) of the injection?Can mod...

Read full article →

Related Articles

NASA’s Mars Sample Return mission is dead
Muhammad523 · Hacker News · 18h ago
AMD's random number generator can't generate a 0?
BruceEel · Hacker News · 5h ago
What happened to the Snowden archive
EXHades · Hacker News · 1d ago
Samsung is expected to more than double output of its HBM4 and HBM4E DRAM
giuliomagnifico · Hacker News · 1d ago
MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis
theanonymousone · Hacker News · 10h ago