OpenVINO Toolkit
OpenVINO is an open-source toolkit from Intel for optimising and deploying AI inference, converting models from frameworks such as PyTorch, TensorFlow and ONNX to run efficiently on Intel CPUs, integrated and discrete GPUs, and NPUs.
Models are converted into OpenVINO's intermediate representation and executed by the OpenVINO Runtime, which provides plugins for different Intel devices and can distribute work across them. Optimisation tools, including the Neural Network Compression Framework, apply quantisation and other compression techniques. OpenVINO Model Server serves models over network APIs, and the toolkit supports computer vision, speech, natural language and generative AI models, including LLMs through dedicated generative AI APIs.
OpenVINO is used to run inference on Intel-based industrial PCs, edge servers and AI PCs for applications such as visual inspection, video analytics, robotics perception and on-device assistants, often without a discrete GPU.
Performance depends on the target device, model architecture, precision and the optimisations applied, so models should be benchmarked on the intended hardware with the tools included in the toolkit. Supported operations, devices and model types change between releases, so the current release notes should be checked.
Key points
- Open-source Intel toolkit for optimising and deploying inference
- Runs models on Intel CPUs, GPUs and NPUs
- Imports models from PyTorch, TensorFlow and ONNX
- Includes compression tools and a model server
Where AiVibe comes in
AiVibe builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.