BeamGPT: A new paradigm for attention

·LessWrong··

I have found an operator that achieves striking results in learning curves when used alongside standard attention in a nanoGPT-style character-level language model. It finds structure in the sequence that attention misses.The model learns a mix ratio of around 45% attention to 55% of the field operator. This ratio seems consistent across layers. This operator is linear in sequence length. Standard attention is quadratic. The hybrid scaling model gives roughly 2.3 savings at long context. As you ...

Read full article →

Related Articles

England set to be one of the first countries to eliminate hepatitis C
stevekemp · Hacker News · 23h ago
London Underground begins scanning passengers' faces
BlueBerry2001 · Hacker News · 1d ago
Beef and dairy drive 41% of biodiversity damage linked to global farmland
robtherobber · Hacker News · 2h ago
Stealing Reasoning Traces from Proprietary LLM APIs
quantumgarbage · Hacker News · 22h ago
CFTC declares market emergency, orders Kalshi to continue to operate in New York
michaefe · Hacker News · 11h ago