When people talk about tiny machine learning, they usually mean image classification. Cameras are everywhere and the demos are visual, so it makes sense. But some of the most successful deployments of tiny models have nothing to do with images. They run on accelerometers, gyroscopes, temperature sensors, and microphones, classifying short windows of time series data in real time. The design constraints here are different enough that lessons from vision do not always transfer.
The first difference is input dimensionality. A small image might be 32 by 32 by 3, which is over three thousand values. A typical sensor window might be 128 samples across three axes, which is under four hundred. That changes the calculus for what kind of architecture makes sense. Deep convolutional stacks that work well on images often have too many parameters relative to the input for sensor tasks, and you can frequently get better results with a much shallower model or even a small recurrent network.
With sensor data, a large fraction of your model's job can be done by good preprocessing. Normalizing per axis using running statistics, removing gravity from accelerometer readings, and computing simple derived features like magnitude or jerk often improves accuracy more than swapping architectures. These steps are cheap, deterministic, and easy to verify, which makes them perfect for a microcontroller. Do them in fixed point if you can, and keep them consistent between training and deployment.
Windowing is another decision that matters more than people expect. How long is each window, and how much do consecutive windows overlap? Shorter windows reduce latency but give the model less context. Longer windows improve accuracy but delay the output. For gesture recognition, windows of one to two seconds with fifty percent overlap are a reasonable starting point, but you should tune this against your actual use case. Remember that overlapping windows mean you run inference more often, which directly affects power draw.
Architecturally, one-dimensional convolutions are a strong default. They capture local patterns along the time axis, they are cheap, and they map cleanly onto the kind of math microcontrollers do well. A few stacked 1D conv layers followed by global average pooling and a small dense head can match much larger models on many sensor tasks. If your data has long-range dependencies, a small gated recurrent unit can help, but recurrent layers are harder to quantize and harder to run efficiently on edge hardware, so reach for them only when the simpler option falls short.
Sensor models are especially prone to a specific failure: they look great on a random train-test split and fall apart in the real world. The reason is usually that adjacent windows from the same recording leak across the split, so the test set is not actually independent. Split your data by recording session or by subject, not by window. This is the single most important evaluation habit for time series, and it is the one most often skipped.
Also test on data collected by a different person, a different device, or a different mounting orientation than your training data. Real deployments face all of these shifts. A model that holds up under those conditions is worth far more than one that squeezes out another point of accuracy on a leaky benchmark.
Tiny models and sensor data are a natural fit. Treat the preprocessing seriously, keep the architecture simple, and evaluate honestly, and you will end up with something that actually works outside the lab.
This incident opened my eyes to the value of Insurance in general, so I decided to examine my personal and business insurance.
Comment (3)
The payments are made directly from one person to another without passing through a central bank or clearing house.
22 feb,2025
ReplyChoosing a Lightweight Framework: TFLite, ONNX Runtime, NCNN, and ExecuTorch Compared
23 feb,2025
ReplyPruning Neural Networks by Hand: A Gentle Introduction to Sparsity
23 feb,2025
Reply