AiVibe

AI & Machine Learning

AI Inference

AI inference is the stage at which a trained model is used to make predictions or generate outputs on new data, as distinct from training, and it determines the latency, throughput and running cost of an AI application.

During inference, input data passes forward through the model with fixed weights, and no learning takes place. For LLMs, inference has two phases: processing the prompt, often called prefill, and generating output tokens one at a time, called decoding, which is usually limited by memory bandwidth. Serving software batches requests, caches intermediate attention data and schedules work across accelerators.

Inference runs in cloud data centres for large models and many users, on premises where data must stay local, and on edge devices for low latency and offline operation, such as inspection cameras on a production line. For widely used models, cumulative inference compute can exceed the compute used to train them.

Key metrics are latency, throughput, time to first token and tokens per second for language models, cost per request and energy use. Optimisations include quantisation, distillation, pruning, batching, compilation with tools such as ONNX Runtime, TensorRT or OpenVINO, and choosing hardware suited to the model.

Key points

Where AiVibe comes in

AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.

Explore AiVibe’s work in AI & Machine Learning →

Related terms

Terms that refer to AI Inference

Ask AiMuruga can explain AI Inference for your plant, product or security programme, and draw how it fits.