Why did it get Sparser?
I think that Polysemanticity in artificial neural networks could be the key to making better and smaller models. I having been working on a little project on trying to induce Polysemanticity at a small scale to compare performance, my first approach was to make the bias more adaptable, I called this the Flexbias, however I got a sparser neural network.What is the flex bias?I used a standard transformer architecture including the MLP, Since i wanted to change how the information is processed I al...
Read full article →