B200 Attention Kernel from Scratch to Near-SOTA in 60 Diagrams
magoghm
22 points
3 comments
September 02, 2026
Related Discussions
Found 5 related stories in 69.7ms across 6,607 title embeddings via pgvector HNSW
- Show HN: LLM Attention Visualization ifz · 150 pts · September 08, 2026 · 48% similar
- Show HN: We built the smallest dual-band aircraft tracker CoolNamesAllTkn · 27 pts · August 26, 2026 · 43% similar
- A walk through of the DeltaNet family of linear attention variants AnhTho_FR · 288 pts · July 28, 2026 · 42% similar
- Understanding FlashAttention Pt 1: Personal Notes ibobev · 13 pts · September 14, 2026 · 42% similar
- Superhuman Attention tsenart · 19 pts · August 28, 2026 · 42% similar
Discussion Highlights (3 comments)
xiphias2
Pretty cool, it's missing how these concepts map to ThunderKittens code (my favorite CUDA toolkit)
erichocean
Really interesting project, appreciated.
peter_d_sherman
One of the best AI GPU Attention Kernel how-to-implement articles I've read in a long time! Tremendous work, tremendous effort was placed into writing this article -- and it shows! Well done!