Challenge: Hand coding weights for efficient sequence memorisation

·LessWrong··

We hand coded weights for one layer MLPs that memorises labels for input token sequences of length two. The number of facts our hand-coded models can memorise with 90% accuracy[1]scales roughly linearly with the models' parameter count[2], just like trained models for the same architecture. However, our hand-coded models' scaling prefactor still falls short of trained models' by a factor of mjx-math { display: inline-block; text-align: left; line-height: 0; text-indent: 0; font-style: normal; fo...

Read full article →

Related Articles

Alphabet's cash burn raises alarm for Big Tech as AI spending climbs
1vuio0pswjnm7 · Hacker News · 7h ago
DARPA, U.S. Air Force fly AI-controlled F-16
r2sk5t · Hacker News · 7h ago
Everyone should know SIMD
WadeGrimridge · Hacker News · 1d ago
Cruller: Bun's Zig Runtime, Continued on Zig 0.16
Erenay09 · Hacker News · 15h ago
LG to ban residential proxies from smart TV apps
DemiGuru · Hacker News · 1d ago