When (and when not) LLMs can verbalize awareness of J-Space concept injections - Initial results

·LessWrong··

Code for reproduction and cross-model extensions available here.SummaryI injected single-token Jacobian Lens (J-Lens) vectors into Qwen 3.6–27B while it answered 20 simple factual questions. The injected concept was either a wrong but task-related answer (for example, Athens while asking for the capital of Egypt) or a wholly unrelated concept. Injections were performed at several strengths over three layer-band categories: the full estimated workspace band, its first half, or its second half.I r...

Read full article →

Related Articles

England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 17h ago
Stealing Reasoning Traces from Proprietary LLM APIs
quantumgarbage · Hacker News · 16h ago
CFTC declares market emergency, orders Kalshi to continue to operate in New York
michaefe · Hacker News · 5h ago
London Underground begins scanning passengers' faces
BlueBerry2001 · Hacker News · 20h ago
Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 1d ago