Modern GPU Programming for MLSys

·Hacker News··

Skip to main content Back to top Ctrl+K Search Ctrl+K Part I, Understanding the GPU GPU Execution Model What Makes a Kernel Fast Data Layout and Its Notation Tensor Core Operand Layouts Across GPU Generations Async Data Movement: TMA Tensor Cores: tcgen05 Special Memory: TMEM Async Coordination: mbarriers Advanced: Cluster Launch Control Part II, TIRx Overview Introduction to TIRx TIRx Layout API Part III, GEMM: Tiled to SOTA Building a Tiled GEMM Pipelining GEMM with TMA Scaling GEMM with Warp

Read full article →

Related Articles

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
riordan · Hacker News · 12h ago
Mistral Patent for “Code implemented tool calls”
theanonymousone · Hacker News · 9h ago
Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints
kotaKat · Hacker News · 7h ago
Study links GLP-1 drugs to bigger jump in women's employment than a degree
metadat · Hacker News · 6h ago
Tail-call optimization in C is relatively recent (2025)
prakashqwerty · Hacker News · 11h ago