B200 Attention Kernel from Scratch to Near-SOTA in 60 Diagrams
magoghm
22 points
3 comments
September 02, 2026
Related Discussions
Found 5 related stories in 53.2ms across 5,346 title embeddings via pgvector HNSW
- Show HN: We built the smallest dual-band aircraft tracker CoolNamesAllTkn · 27 pts · August 26, 2026 · 43% similar
- A walk through of the DeltaNet family of linear attention variants AnhTho_FR · 288 pts · July 28, 2026 · 42% similar
- Superhuman Attention tsenart · 19 pts · August 28, 2026 · 42% similar
- Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels rawsh · 33 pts · July 12, 2026 · 41% similar
- OE3GBB tehsauce · 13 pts · July 21, 2026 · 41% similar
Discussion Highlights (3 comments)
xiphias2
Pretty cool, it's missing how these concepts map to ThunderKittens code (my favorite CUDA toolkit)
erichocean
Really interesting project, appreciated.
peter_d_sherman
One of the best AI GPU Attention Kernel how-to-implement articles I've read in a long time! Tremendous work, tremendous effort was placed into writing this article -- and it shows! Well done!