Vision-Language Model (VLM)
A vision-language model (VLM) is a multimodal model that processes images together with text, so that it can describe images, answer questions about them, read documents and charts, and follow instructions that refer to visual content.
A typical VLM pairs an image encoder, often a vision transformer, with a large language model. Image features are projected into the language model's embedding space and processed alongside text tokens. Contrastive models such as CLIP instead learn a shared embedding space for images and text, enabling zero-shot classification and image search. Training uses large collections of image-text pairs, followed by instruction tuning.
Applications include visual question answering, document and form understanding, chart reading, image captioning, describing defects in natural language and assistants that interpret photographs of equipment or screens. VLMs are also the starting point for vision-language-action models in robotics.
VLMs can misread fine details, small text and precise measurements, and can describe things that are not present. For measurement-critical inspection, dedicated machine vision with calibrated optics remains the established approach. Image resolution limits, latency and the cost of processing each image as many tokens must be considered.
Key points
- Processes images and text together in one model
- Typically an image encoder connected to a large language model
- CLIP-style models learn a shared image-text embedding space
- Basis for vision-language-action models in robotics
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.
Related terms
- Large Language Model (LLM)AI & Machine Learning
- Computer VisionAI & Machine Learning
- Vision-Language-Action Model (VLA)Robotics & Physical AI
- Transformer (Neural Network Architecture)AI & Machine Learning
- Embeddings (Vector Embeddings)AI & Machine Learning