Reasoning and learning about injected concepts in language models

·LessWrong··

This work was done as a part of SPAR, under the mentorship of Mirko Bronzi and Damiano Fornasiere. TL;DRWe test models' ability to recover information about their activations by injecting steering vectors, and asking the LLMs to verbalize properties of them. We train models with in-context learning and test for three capabilities: Can models identify the region of layers (early, middle, late) of the injection?Can models identify the relative magnitude (low, medium, high) of the injection?Can mod...

Read full article →

Related Articles

US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 22h ago
Oracle bans AI-generated code from OpenJDK
delduca · Hacker News · 15h ago
AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 1d ago
Making Postgres 300x faster for analytics: batching, operator fusion, and SIMD
poly2it · Hacker News · 21h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 1d ago