AiVibe

AI & Machine Learning

Data Annotation (Data Labelling)

Data annotation, or labelling, is the process of adding correct answers to raw data, such as class labels, bounding boxes, segmentation masks or transcriptions, so that the data can be used to train and evaluate supervised machine learning models.

Annotation types depend on the task: image-level labels for classification, bounding boxes for object detection, pixel masks for segmentation, keypoints for pose estimation, text spans for entity extraction and transcripts for speech. Annotators follow written guidelines using labelling tools, and quality is controlled through review, consensus between annotators and agreement metrics. Model-assisted labelling and active learning reduce effort by pre-labelling data and prioritising uncertain examples.

In manufacturing, annotation is often the largest effort in a visual inspection project, because defects must be labelled consistently by people who understand the quality criteria. Defect catalogues and boundary samples agreed with quality teams help keep labels consistent.

Inconsistent or wrong labels limit achievable accuracy and distort evaluation. Guidelines should define edge cases, and test sets should be labelled with extra care. Confidential images and personal data in datasets require access control, and outsourcing labelling raises data protection considerations.

Key points

Where AiVibe comes in

AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.

Explore AiVibe’s work in AI & Machine Learning →

Related terms

Terms that refer to Data Annotation (Data Labelling)

Ask AiMuruga can explain Data Annotation (Data Labelling) for your plant, product or security programme, and draw how it fits.