Model Serving
Model serving is the deployment of trained machine-learning models behind an API or batch process so applications can obtain predictions, with the serving layer handling scaling, batching, versioning and monitoring.
Model serving is the process of deploying a trained machine-learning model so that applications can send it input data and receive predictions. A serving system loads the model into memory on suitable hardware, exposes it through an API, typically HTTP/REST or gRPC, and manages concerns such as request batching, concurrency, model versioning, scaling and monitoring. Serving can be online, returning predictions in real time per request, or batch, scoring large datasets on a schedule.
Open-source serving frameworks include NVIDIA Triton Inference Server, KServe for Kubernetes, TensorFlow Serving and vLLM for large language models, while managed platforms such as Amazon SageMaker, Azure Machine Learning and Google Vertex AI provide hosted endpoints. In manufacturing, models for visual inspection or anomaly detection are often served on edge devices near the line to avoid network latency, with toolkits such as OpenVINO optimising them for Intel hardware.
Key metrics are latency, usually tracked at high percentiles, throughput, error rate, hardware utilisation and cost per prediction. Model outputs should also be monitored for drift and quality problems, which infrastructure metrics do not reveal. Optimisations such as dynamic batching, quantisation and compiling models for specific hardware improve efficiency, and new model versions are commonly rolled out with canary or shadow deployments.
Key points
- Online serving answers individual requests; batch serving scores datasets on a schedule
- Frameworks include NVIDIA Triton, KServe, TensorFlow Serving and vLLM
- Latency percentiles, throughput, utilisation and cost per prediction are key metrics
- Edge serving keeps latency-sensitive inference, such as visual inspection, near the line
Where AiVibe comes in
AiVibe builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.