YFarmX

Flash attention

AI

Flash attention: An input-output-aware algorithm that computes exact attention without storing the full attention matrix in memory, greatly speeding training and inference.

Related terms

Browse the full glossary →