Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

·Hacker News··

I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift.It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next

Read full article →

Related Articles

The ChatGPT/Codex app bundles a full copy of LibreOffice
timpera · Hacker News · 13h ago
I trained a small transformer in 1.5hrs and it beats many LLMs
porridgeraisin · Hacker News · 23h ago
The Emergent Symbolic Structure of Artificial Neural Networks
schmuhblaster · Hacker News · 5h ago
Refurbishing a Tektronix TDS7104 Oscilloscope
jwise0 · Hacker News · 13h ago
Omarchy: Any User Process Can Escalate to Root
trap0xcc · Hacker News · 2d ago