Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

·Hacker News··

How I built a 5.01 BPW Qwen3.8 27B hybrid, filled its 256K context, tuned MTP and gained 21.97% with a measured llama.cpp runtime.

Read full article →

Related Articles

Self hosted email continues to steeply decline
minusf · Hacker News · 6h ago
Apple's App Tracking Transparency treated its own apps better than rivals
nyku · Hacker News · 3h ago
AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira
galnagli · Hacker News · 3h ago
AI has access to a vastly larger working memory than the human brain
rzk · Hacker News · 1d ago
A Preview of DuckDB v2.0
ibotty · Hacker News · 3h ago