TOPS (Tera Operations Per Second)
TOPS, or tera operations per second, is a measure of an AI accelerator's peak compute throughput: one TOPS equals one trillion operations per second, and figures are typically quoted for low-precision integer arithmetic such as INT8.
Vendors calculate TOPS from the number of multiply-accumulate units, the clock frequency and the operations per cycle, usually counting each multiply-accumulate as two operations. The figure is a theoretical peak, often stated at INT8 or lower precision, and sometimes assumes sparsity features that skip zero values. TFLOPS is the equivalent measure for floating-point operations.
TOPS figures are widely used to compare NPUs, edge AI modules and AI PCs, and to give a rough sense of whether a device can handle a workload such as several camera streams running object detection.
Real performance depends on memory bandwidth, supported operations, software maturity, model architecture and how well the model maps onto the hardware, so devices with similar TOPS can perform very differently. Comparisons should state the precision and whether sparsity is assumed, and decisions should rest on benchmarks of the actual model, measuring frames or tokens per second, latency and power. MLPerf benchmarks from MLCommons provide standardised comparisons.
Key points
- One TOPS is one trillion operations per second
- A theoretical peak, usually quoted at INT8 precision
- Precision and sparsity assumptions must be stated for comparisons
- Benchmarks of the real model matter more than peak TOPS
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.