Challenge: Hand coding weights for efficient sequence memorisation

·LessWrong··

We hand coded weights for one layer MLPs that memorises labels for input token sequences of length two. The number of facts our hand-coded models can memorise with 90% accuracy[1]scales roughly linearly with the models' parameter count[2], just like trained models for the same architecture. However, our hand-coded models' scaling prefactor still falls short of trained models' by a factor of mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; fo...

Read full article →

Related Articles

Asahi Linux Now Officially Supports Apple M3 Macs – With Caveats
mdp2021 · Hacker News · 7h ago
Private German rocket makes history, reaches orbit from European soil
bookmtn · Hacker News · 1d ago
LLMs as a Cognitive Virus
canjobear · Hacker News · 1d ago
Actively exploited sandbox RCE in all Chromium versions
negura · Hacker News · 1d ago
The car industry A/B tested selling a car with and without CarPlay
gumby · Hacker News · 1h ago