Exploring Generalization in NLA's

·LessWrong··

Recently, I was reading anthropic's paper on NLA's[1] and for a person who works on steering, it was an interesting and thought-provoking paper. In this post I would like to go through my reproduction and some of the experiments I did on them.Training and ArchitectureI'm going to touch little on architecture here because the paper already covers them, I add it here so that it could make little sense or give a refresh while reading. So, we basically train 2 models,Activation Verbalizer (AV): Inje...

Read full article →

Related Articles

Claude Opus 5.5
km144 · Hacker News · 1d ago
GPT-6 Astra has gained the ability to drive a car
plurby · Hacker News · 15h ago
Claude Code reads AGENTS.md only when telemetry is on [fixed]
pszypowicz · Hacker News · 18h ago
UK military jamming other nations' satellites to defend itself, BBC told
thm · Hacker News · 12h ago
What California is learning from solar panels built over irrigation canals
Jtsummers · Hacker News · 2d ago