Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

·Hacker News··

How I built a 5.01 BPW Qwen3.8 27B hybrid, filled its 256K context, tuned MTP and gained 21.97% with a measured llama.cpp runtime.

Read full article →

Related Articles

Pi 1.0
sergiotapia · Hacker News · 6h ago
Automatic Transmission – a data-privacy study of connected vehicles
rafaelc · Hacker News · 5h ago
Cops Can Bypass iPhone's Automatic Reboot to Get into Locked Phones
speckx · Hacker News · 11h ago
Cloudflare K2: serverless event streams
elffjs · Hacker News · 11h ago
Singapore govt dating app uses Gale-Shapley stable marriage algorithm
rzk · Hacker News · 1d ago