Running Kimi K3 on a M1 Max

·Hacker News··

Run Kimi K3, a 2.8T-parameter Mixture-of-Experts LLM, on a single Apple Silicon Mac. Streams MXFP4 experts on demand over HTTP into a local disk cache — fused NEON kernels, Metal/MPS compute, exact reproducible decoding, and an OpenAI-compatible API server for local chat and coding agents. - gavamedia/deltafin

Read full article →

Related Articles

New HIV vaccine shows unprecedented success in preclinical study
codebyaditya · Hacker News · 11h ago
Zig's Incremental Compilation Internals
garyhtou · Hacker News · 8h ago
Discovering Cryptographic Weaknesses with Claude
gslin · Hacker News · 7h ago
A walk through of the DeltaNet family of linear attention variants
AnhTho_FR · Hacker News · 8h ago
US citizen charged after GrapheneOS phone wipes during airport search
eecc · Hacker News · 2d ago