GateGPT: 56k tokens per second Transformer (KV cache) on FPGA at 80 MHz

·Hacker News··

56,000+ tokens/sec at just 80 MHz. 🤯 I burned a full Transformer with KV cache into a custom chip. Designed gate by gate as a 100% digital integrated circuit. Prototyped on a FPGA. (No GPU. No CPU) Just pure digital silicon running @karpathy microGPT, spelling out names on a https://t.co/hXX8kKIxiA

Read full article →

Related Articles

Show HN: Shitty – fast terminal. Memory-unsafe and faster than yours
pshirshov · Hacker News · 10h ago
EU Age Verification Project Mandates Hardware-Bound Attestation
RobotToaster · Hacker News · 13h ago
AI migrated legacy COBOL programs to Java, bugs included
felineflock · Hacker News · 6h ago
Go 1.27 Interactive Tour
Hixon10 · Hacker News · 1d ago
Google fixed more Chrome bugs in June than over the past two years, thanks to AI
Garbage · Hacker News · 3d ago