Open Neural Network Exchange (ONNX)
ONNX (Open Neural Network Exchange) is an open format for representing machine learning models, which lets a model trained in one framework, such as PyTorch, be exported and run with different runtimes and on different hardware.
An ONNX file describes the model as a computation graph of standard operators, defined in versioned operator sets, together with the trained weights. Frameworks export to ONNX, and runtimes such as ONNX Runtime execute it, using execution providers that target CPUs, GPUs and NPUs through back ends such as CUDA, TensorRT, DirectML and OpenVINO. The project was started by Microsoft and Facebook in 2017 and is now part of the Linux Foundation's AI and Data foundation.
ONNX decouples training from deployment: data scientists train in their preferred framework, and engineers deploy the same model to servers, desktops or edge devices with an optimised runtime. Many hardware vendors' toolkits accept ONNX models as input.
Not every framework operation exports cleanly; custom layers or newer operators may need workarounds, and outputs should be compared numerically between the original and exported models. The operator set version must be supported by the target runtime. ONNX models can also be quantised for faster inference.
Key points
- An open, framework-neutral format for ML models
- Represents models as graphs of standard operators plus weights
- ONNX Runtime executes models on many hardware back ends
- Exported outputs should be checked against the original model
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.