MoonshotAI/FlashKDA
FlashKDA: high-performance Kimi Delta Attention kernelsFlashKDA FlashKDA: Flash Kimi Delta Attention — high-performance KDA kernels built on CUTLASS News 2026-04-22 — Deep-Dive Blog: the design decisions behind FlashKDA v1, read it here. Requirements SM90 and above CUDA 12.9 and above PyTorch 2.4 and above Installation git clone https://github.com/MoonshotAI/FlashKDA.git flash-kda cd flash-kda git submodule update --init --recursive pip install -v --no-build-isolation . By default, the build det...
Read full article →