Model Evaluation
Model evaluation measures how well a machine learning model performs on data it has not seen during training, using metrics suited to the task, in order to decide whether it is fit for deployment and to compare alternatives.
Data is split into training, validation and test sets, or evaluated with cross-validation. Classification models are assessed with a confusion matrix and metrics such as accuracy, precision, recall, F1 score and ROC-AUC; regression models with mean absolute error or root mean squared error; and object detectors with mean average precision. Decision thresholds trade false positives against false negatives.
In visual inspection, missed defects and false rejects have direct cost implications, and in predictive maintenance so do missed failures and false alarms. Evaluation should therefore use metrics tied to business and safety consequences, and test data should represent real operating conditions, including different shifts, product variants and seasons.
Accuracy alone misleads on imbalanced data such as rare defects. Beyond aggregate metrics, evaluation examines performance across subgroups and conditions, the calibration of confidence scores, robustness and fairness. Evaluation continues after deployment through monitoring, since performance can degrade as data drifts.
Key points
- Measures performance on data not used for training
- Metrics include precision, recall, F1, ROC-AUC, MAE and mAP
- Accuracy alone misleads when classes are imbalanced
- Metrics should reflect the real cost of each error type
Where AiVibe comes in
AiVibe's AI quality management services include bias detection and model validation.
Related terms
- Overfitting and UnderfittingAI & Machine Learning
- Supervised LearningAI & Machine Learning
- LLM EvaluationAI & Machine Learning
- Model Drift (Data and Concept Drift)AI & Machine Learning
- AI BiasAI & Machine Learning
- Anomaly DetectionAI & Machine Learning