YFarmX

Vision-language model (VLM)

AI

Vision-language model (VLM): A model that takes images and text together and reasons across both, the architecture behind screenshot understanding and document extraction.

Related terms

Browse the full glossary →