tracked author
Tri Dao
FlashAttention first author and FlashAttention-2 single author; current public profile connects Princeton, Together AI, and the FlashAttention implementation line.
Representative Papers
来自作者已核验个人主页的重点论文;本站单篇归档见下方 Related Notes。
- 01 Marconi: Prefix Caching for the Era of Hybrid LLMs MLSys · 2025
- 02 FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision NeurIPS · 2024
- 03 Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality ICML · 2024
- 04 Mamba: Linear-Time Sequence Modeling with Selective State Spaces COLM · 2024
- 05 FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness NeurIPS · 2022
- 06 Monarch: Expressive Structured Matrices for Efficient and Accurate Training ICML · 2022
Related Notes
按论文归档时间排序,展示该作者在本站已经出现的材料。
归档 更新
FlashAttention 2: Faster Attention with Better Parallelism and Work Partitioning
通过减少 non matmul FLOPs、提升 sequence parallelism 和调整 warp 分工提高 attention kernel 利用率。
归档 更新
FlashAttention: Fast and Memory Efficient Exact Attention with IO Awareness
用 SRAM tiling、online softmax 与 backward recomputation 减少 exact attention 的 HBM traffic。