MoonshotAI/FlashKDA

GitHub Trending··

FlashKDA: high-performance Kimi Delta Attention kernelsFlashKDA FlashKDA: Flash Kimi Delta Attention — high-performance KDA kernels built on CUTLASS News 2026-04-22 — Deep-Dive Blog: the design decisions behind FlashKDA v1, read it here. Requirements SM90 and above CUDA 12.9 and above PyTorch 2.4 and above Installation git clone https://github.com/MoonshotAI/FlashKDA.git flash-kda cd flash-kda git submodule update --init --recursive pip install -v --no-build-isolation . By default, the build det...

Read full article →

Related Articles

OpenAI’s o1 correctly diagnosed 67% of ER patients vs. 50-55% by triage doctors
donsupreme · Hacker News · 2mo ago
Accelerating Gemma 4: faster inference with multi-token prediction drafters
amrrs · Hacker News · 2mo ago
A couple million lines of Haskell: Production engineering at Mercury
unignorant · Hacker News · 2mo ago
Using “underdrawings” for accurate text and numbers
samcollins · Hacker News · 2mo ago
ProgramBench: Can language models rebuild programs from scratch?
jonbaer · Hacker News · 2mo ago