Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

·Hacker News··

How I built a 5.01 BPW Qwen3.8 27B hybrid, filled its 256K context, tuned MTP and gained 21.97% with a measured llama.cpp runtime.

Read full article →

Related Articles

LG smart TVs caught logging audio with screen off and snooping on local devices
chris_overseas · Hacker News · 14h ago
Smartphone makers don't bother to comply with EU repairability requirements
mdp2021 · Hacker News · 10h ago
bzip3
tosh · Hacker News · 8h ago
Asahi Linux on M3
mdp2021 · Hacker News · 1d ago
It took a year to ship WebAssembly in Anubis
xena · Hacker News · 1d ago