Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP
How I built a 5.01 BPW Qwen3.8 27B hybrid, filled its 256K context, tuned MTP and gained 21.97% with a measured llama.cpp runtime.
Read full article →How I built a 5.01 BPW Qwen3.8 27B hybrid, filled its 256K context, tuned MTP and gained 21.97% with a measured llama.cpp runtime.
Read full article →