tracked author
Tri Dao
FlashAttention first author and FlashAttention-2 single author; current public profile connects Princeton, Together AI, and the FlashAttention implementation line.
Related Notes
按论文归档时间排序,展示该作者在本站已经出现的材料。
Muon: An optimizer for hidden layers in neural networks
Muon 的核心技术含量在于把隐藏层矩阵参数的 momentum update 映射到近似半正交方向:它用低精度 Newton Schulz 近似 polar factor,让更新能量更均匀地覆盖矩阵奇异方向;后续大模型实践表明,Muon 想稳定扩展到 LLM pretraining,还需要 weight decay、shape aware update scale、AdamW/Muon 参数分组、分布式 full matrix or...
FlashAttention 2: Faster Attention with Better Parallelism and Work Partitioning
FlashAttention 2 的核心贡献是把 FlashAttention v1 解决 HBM traffic 后剩下的性能瓶颈继续拆开:减少 expensive non matmul FLOPs,把 attention 计算沿 sequence dimension 分给更多 thread blocks 提高 SM occupancy,并把 warp 内 work partition 从 sliced K 调整为 sliced Q...
FlashAttention: Fast and Memory Efficient Exact Attention with IO Awareness
FlashAttention 的核心贡献是把 exact softmax attention 的瓶颈从 FLOPs 视角重新定位到 GPU memory hierarchy 和 HBM 读写上:它用 tiling 在 SRAM 中分块计算 attention,并用 online softmax 统计量与 backward recomputation 避免物化 $N\times N$ attention matrix,从而保持 exac...