Model Size Scaling in 2023-2031

·LessWrong··

Token generation speed is constrained by the speed at which the relevant HBM can be read, which is mostly the weights and KV-cache. Suppose a model is large, so that more than half of HBM is read when making a single pass over the weights, it's being read in parallel within a scale-up system, and N such systems are used in a pipeline. Then the time it takes to generate a token (without speculative decoding) is at least the time of reading more than half of an HBM stack times N. If we target a pa...

Read full article →

Related Articles

AMD acquires Taalas to boost inference performance by etching models in silicon
itvision · Hacker News · 10h ago
Qwen3.8 Max now ranked as the best overall model by agentic index
apitman · Hacker News · 12h ago
Nashville uses eminent domain to block data center near zoo
mapping365 · Hacker News · 1d ago
Launch HN: ProvenMetal (YC S26) delivers circuit boards in days instead of weeks
willcarkner · Hacker News · 15h ago
Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 2d ago