Exploring Generalization in NLA's

·LessWrong··

Recently, I was reading anthropic's paper on NLA's[1] and for a person who works on steering, it was an interesting and thought-provoking paper. In this post I would like to go through my reproduction and some of the experiments I did on them.Training and ArchitectureI'm going to touch little on architecture here because the paper already covers them, I add it here so that it could make little sense or give a refresh while reading. So, we basically train 2 models,Activation Verbalizer (AV): Inje...

Read full article →

Related Articles

Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 1d ago
We replaced Redis with MySQL for inventory reservations and it scaled
adletbalzhanov · Hacker News · 1d ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 2d ago
FCC moves to ban Lidar-equipped foreign drones from US
f-serif · Hacker News · 11h ago
Tom Stanton's supersonic trebuchet breaks sound barrier with gravity alone
Thorondor · Hacker News · 12h ago