YFarmX

Vision transformer (ViT)

AI

Vision transformer (ViT): A model that splits an image into patches, treats them as tokens, and applies transformer attention for image understanding.

Related terms

Browse the full glossary →