YFarmX

FlashAttention

AI

FlashAttention: An optimised way of computing attention that avoids writing large intermediate matrices to memory, making inference and training faster and more memory-efficient.

Related terms

Browse the full glossary →