Exploration: fine-tuning with parameter decomposition

·LessWrong··

TL;DR: We can destroy a 67M-parameter language model's ability to predict German text by fine-tuning a single number: the scalar prefactor on one German-related rank-1 parameter subcomponent. This is an early exploration into using parameter decomposition for a more targeted and interpretable form of model fine-tuning. At small German-token budgets, fine-tuning the scalar prefactor of a single German-related parameter subcomponent beats rank-1 and rank-4 LoRA [1] fine-tunes on the trade-off betw...

Read full article →

Related Articles

We replaced Redis with MySQL for inventory reservations and it scaled
adletbalzhanov · Hacker News · 22h ago
Timeline of the OpenAI accidental attack against Hugging Face
882542F3884314B · Hacker News · 1d ago
Melatonin impairs morning cognition in healthy young adults (2023)
bohaska · Hacker News · 19h ago
FCC moves to ban Lidar-equipped foreign drones from US
f-serif · Hacker News · 4h ago
US strikes $1.2B deal to pay German firm to halt offshore wind projects
defrost · Hacker News · 2d ago