Attention sink
AIAttention sink: The tendency of attention to dump spare weight onto the first few tokens, which quantisation and cache-eviction schemes must preserve.
Attention sink: The tendency of attention to dump spare weight onto the first few tokens, which quantisation and cache-eviction schemes must preserve.