Inference Endpoint
An inference endpoint is a hosted, network-accessible interface, usually an HTTPS API, that receives input data, runs it through a deployed machine-learning model and returns predictions or generated output.
An inference endpoint is a network-addressable interface, usually an HTTPS URL, through which applications send input data to a deployed machine-learning model and receive its predictions or generated output. Behind the endpoint, the hosting platform runs one or more model replicas on CPUs, GPUs or other accelerators, routes requests among them and scales capacity according to traffic. Endpoints are typically secured with API keys, tokens or cloud identity policies.
Managed endpoints are offered by machine-learning platforms such as Amazon SageMaker, Azure Machine Learning and Google Vertex AI, and model providers expose large language models through hosted APIs. Variants include real-time endpoints for low-latency requests, serverless endpoints that scale to zero between requests, asynchronous endpoints for long-running jobs and batch processing for offline scoring. The choice depends on latency needs, traffic patterns and cost.
Important considerations include cold-start delays when capacity scales up from zero, request payload and timeout limits, rate limits and the cost of keeping GPU-backed replicas running. Inputs and outputs may contain sensitive data, so endpoints should use private networking where possible, encrypt traffic and log requests in line with data protection rules. Monitoring should cover latency, errors, utilisation and the quality of the model's outputs.
Key points
- Usually an HTTPS API backed by one or more model replicas on CPUs or GPUs
- Types include real-time, serverless, asynchronous and batch inference
- Secured with API keys, tokens or cloud identity policies and private networking
- Cold starts, payload limits and GPU running costs shape endpoint design
Where AiVibe comes in
AiVibe Software Services delivers cloud solutions on AWS, Microsoft Azure, Google Cloud or on-premise, together with cloud security, legacy modernisation, data analytics and AI and machine learning services.