Integer Quantization: Deep Dive

·Hacker News··

A lot has happened in transformer quantization over the past few years, from barely being able to quantize a 7B model in INT8 without...

Read full article →

Related Articles

Ten advances in mathematics and theoretical computer science
milkshakes · Hacker News · 17h ago
Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
leonickson · Hacker News · 16h ago
MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
vblanco · Hacker News · 20h ago
AirLLM 70B inference with single 4GB GPU
Anon84 · Hacker News · 22h ago
Smaller, faster, safer: running Kimi and GLM at scale
ascorbic · Hacker News · 16h ago