Smaller, faster, safer: running Kimi and GLM at scale

·Hacker News··

Serving frontier models like Kimi and GLM means fighting for GPU memory. Here's how we quantize KV caches, compress model weights, and add integrity checks to serve them faster, cheaper, and safely.

Read full article →

Related Articles

Ten advances in mathematics and theoretical computer science
milkshakes · Hacker News · 3h ago
MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
vblanco · Hacker News · 5h ago
Rust project goals: Immobile types and guaranteed destructors
paavohtl · Hacker News · 12h ago
AirLLM 70B inference with single 4GB GPU
Anon84 · Hacker News · 8h ago
Show HN: Shitty – fast terminal. Memory-unsafe and faster than yours
pshirshov · Hacker News · 20h ago