Model Quantisation
Model quantisation reduces the numerical precision of a neural network's weights and activations, for example from 32-bit floating point to 8-bit or 4-bit integers, making models smaller and faster, usually with limited loss of accuracy.
Quantisation maps a range of floating-point values onto a small set of integers using scale factors. Post-training quantisation converts an already trained model, often using a small calibration dataset to choose value ranges, while quantisation-aware training simulates low precision during training to preserve accuracy. Large language models are commonly quantised to 8-bit or 4-bit weights with methods such as GPTQ and AWQ, and formats such as GGUF package quantised models for local inference.
Quantisation lets models run on edge devices, CPUs and NPUs with integer acceleration, reduces memory so larger models fit on available GPUs, and increases throughput while lowering energy use. INT8 inference is widely supported by edge AI toolkits.
Accuracy loss varies by model and task and must be measured on representative validation data, and sensitive layers can be kept at higher precision. Speed gains depend on hardware support for the chosen format. Documentation and tools generally use the American spelling, quantization.
Key points
- Lowers the numerical precision of weights and activations
- Post-training quantisation and quantisation-aware training are the main approaches
- Shrinks models and speeds up inference on edge hardware
- Accuracy impact must be measured on validation data
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.
Related terms
- Knowledge DistillationAI & Machine Learning
- Edge AIAI & Machine Learning
- AI InferenceAI & Machine Learning
- Neural Processing Unit (NPU)AI & Machine Learning
- OpenVINO ToolkitAI & Machine Learning
- Fine-TuningAI & Machine Learning