Image Segmentation
Image segmentation is the computer vision task of classifying every pixel in an image, producing masks that outline the exact shape and area of objects, defects or regions rather than just boxes or a single label.
Semantic segmentation assigns a class to each pixel without separating individual objects, instance segmentation separates each object, and panoptic segmentation combines both. Architectures include U-Net, originally developed for biomedical images, DeepLab and Mask R-CNN, which extends a detector with a mask branch. Promptable foundation models such as Meta's Segment Anything Model can segment objects indicated by points or boxes without task-specific training.
Segmentation measures defect size and area for grading, inspects coatings, welds and seals, separates overlapping parts for robot picking, analyses textile and surface textures, and supports medical and satellite image analysis. It is preferred when a pass or fail decision depends on the size or shape of a defect.
Pixel-level annotation is labour-intensive, although model-assisted tools reduce the effort. Performance is measured with intersection over union and the Dice coefficient. Fine boundaries and thin defects such as hairline cracks demand sufficient image resolution, and inference is typically heavier than for classification.
Key points
- Classifies each pixel to outline objects and defects precisely
- Semantic, instance and panoptic variants exist
- U-Net and Mask R-CNN are widely used architectures
- Measured with intersection over union and the Dice coefficient
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.