Quantization is the process of representing model weights and activations with fewer bits than the float32 they were trained with. The most common target is int8, which uses 8 bits per value instead of 32. That is a 4x reduction in model size, and on hardware with integer arithmetic it can also be a significant speedup. The catch is that you lose precision, and if you do it carelessly, accuracy drops.
The good news is that modern quantization workflows are much less painful than they used to be. You do not need to retrain from scratch in most cases. You need to understand what is happening and follow a few steps in the right order.
Post-training quantization (PTQ) takes a trained float model and converts it to int8 without any further training. It requires a small representative dataset, usually a few hundred samples, to calibrate the ranges of activations. This is the easiest path and often works well for convolutional networks. If your accuracy drop is acceptable, you are done.
When PTQ hurts too much, quantization-aware training (QAT) is the next step. QAT simulates the effects of quantization during training by inserting fake quantization nodes into the graph. The model learns to be robust to the reduced precision. This usually recovers most of the lost accuracy, at the cost of a full training run. It is more work, but for models where every point of accuracy matters, it is worth it.
A practical rule: try PTQ first. If the accuracy drop is under a percentage point or two, ship it. If it is larger, look at which layers are most sensitive before jumping to QAT. Often a handful of layers, typically the first and last, benefit from staying in higher precision. Many frameworks let you keep those layers in float while quantizing the rest.
The most common failure mode is activation ranges. If a layer produces occasional very large values, the calibration step may choose a wide range that compresses most values into a small number of int8 buckets. The fix is usually to clip the outliers, either by adjusting the calibration method or by modifying the model to avoid extreme activations. Batch normalization placement matters here; folding it into the preceding convolution before quantization is standard practice.
Another issue is depthwise convolutions. They tend to have per-channel weight distributions that vary a lot, and per-tensor quantization can struggle. Per-channel quantization, where each channel gets its own scale, is the standard solution and is supported in most toolchains. If you are seeing accuracy loss concentrated in depthwise layers, this is the first thing to check.
Finally, do not forget the input. If your model expects normalized float inputs, you need to handle that conversion somewhere. Some frameworks quantize the input tensor as well, which means you feed int8 values directly. Others keep the input in float and quantize internally. Either is fine, but you need to match what the model expects, and getting this wrong produces silently wrong outputs rather than an error.
The overall workflow is: train in float, try PTQ with a good calibration set, measure accuracy on a held-out set, and only move to QAT if you have to. Most edge deployments never need QAT. But knowing it is there, and knowing when to reach for it, is what separates a smooth deployment from a frustrating one.
This incident opened my eyes to the value of Insurance in general, so I decided to examine my personal and business insurance.
Comment (3)
The payments are made directly from one person to another without passing through a central bank or clearing house.
22 feb,2025
ReplyDesigning Tiny Models for Sensor Data: Lessons from Time Series
23 feb,2025
ReplyA Practical Comparison of Lightweight Inference Runtimes for Android
23 feb,2025
Reply