BeamGPT: A new paradigm for attention

·LessWrong··

I have found an operator that achieves striking results in learning curves when used alongside standard attention in a nanoGPT-style character-level language model. It finds structure in the sequence that attention misses.The model learns a mix ratio of around 45% attention to 55% of the field operator. This ratio seems consistent across layers. This operator is linear in sequence length. Standard attention is quadratic. The hybrid scaling model gives roughly 2.3 savings at long context. As you ...

Read full article →

Related Articles

Revealing the details of how OpenAI agents hacked Hugging Face
specked-citrus · Hacker News · 18h ago
Dutch governments builds alternative for Microsoft based on NixOS
fjfaase · Hacker News · 1d ago
Ask HN: Who's still keeping a DOS machine up because the business depends on it?
mlaux · Hacker News · 19h ago
Excel now supports multiple values in a single cell
luispa · Hacker News · 18h ago
Toyota is taking the Corolla electric
cisc · Hacker News · 2d ago