AI Inference
AI inference is the stage at which a trained model is used to make predictions or generate outputs on new data, as distinct from training, and it determines the latency, throughput and running cost of an AI application.
During inference, input data passes forward through the model with fixed weights, and no learning takes place. For LLMs, inference has two phases: processing the prompt, often called prefill, and generating output tokens one at a time, called decoding, which is usually limited by memory bandwidth. Serving software batches requests, caches intermediate attention data and schedules work across accelerators.
Inference runs in cloud data centres for large models and many users, on premises where data must stay local, and on edge devices for low latency and offline operation, such as inspection cameras on a production line. For widely used models, cumulative inference compute can exceed the compute used to train them.
Key metrics are latency, throughput, time to first token and tokens per second for language models, cost per request and energy use. Optimisations include quantisation, distillation, pruning, batching, compilation with tools such as ONNX Runtime, TensorRT or OpenVINO, and choosing hardware suited to the model.
Key points
- Using a trained model on new data, with no learning taking place
- Drives the latency, throughput and running cost of AI applications
- LLM inference has prefill and token-by-token decoding phases
- Optimised through quantisation, batching and compilation
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.
Related terms
- Edge AIAI & Machine Learning
- Model QuantisationAI & Machine Learning
- Graphics Processing Unit (GPU) for AIAI & Machine Learning
- OpenVINO ToolkitAI & Machine Learning
- Open Neural Network Exchange (ONNX)AI & Machine Learning
- Large Language Model (LLM)AI & Machine Learning