Inference cost at scale with napkin math

·Hacker News··

index about blog work now Inference cost at scale with napkin math Jun 14 gpu ai math If you serve AI models as a part of your product stack, you've likely wondered what kind of scale your GPU cluster tops out at. With some rudimentary knowledge about your hardware and model architecture, we can work out the dollar cost-per-user on the back of a napkin1. If you're comfortable reasoning about GPUs and/or LLMs, use this legend to skip to sections of relevance: Resources on a single GPU Cos

Read full article →

Related Articles

Xbox goes down. You can't play games you own on disc
surprisetalk · Hacker News · 13h ago
Ten advances in mathematics and theoretical computer science
milkshakes · Hacker News · 1d ago
Germany Records Historic 12B KWh Solar Feed-In in July 2026
johnbarron · Hacker News · 11h ago
Oxide Computer raises $445M (SEC Form D)
depr · Hacker News · 4h ago
Thanks FedEx, This Is Why We Keep Getting Phished (2024)
stymaar · Hacker News · 3h ago