GPU training optimization
FlashAttention 1–4: How IO-Awareness Reshaped the Attention Kernel
A technical walkthrough of FlashAttention’s four generations — from IO-aware tiling on A100, through better work …
GPU training optimization
A technical walkthrough of FlashAttention’s four generations — from IO-aware tiling on A100, through better work …
GPU training optimization
在 RTX 5090 上实测一步 Transformer 训练的时间和显存,再把 RMSNorm 和 attention 逐个 op 拆开,看每一步算了多少、搬了多少、为反向存了多少。最后都落到 attention 的两个 seq × …