Blog Single style 1

September 16, 2026 Post by : Editorial Edge AI Tutorials
A Practical Comparison of Lightweight Inference Runtimes for Android Running a Keyword Spotting Model on a Battery-Powered Microcontroller

Keyword spotting is the classic first edge AI project, and for good reason. The problem is well defined, the datasets are public, and the model can be tiny. But the moment you move from a laptop demo to a battery-powered device that needs to listen for months, a new set of constraints shows up that no tutorial prepares you for. This article is about those constraints, not about training the model itself.

The first thing to understand is that always-on audio is a duty cycle problem. Your microphone and your feature extractor do not sleep, but your neural network should. A typical design runs a small voice activity detector or a simple energy threshold on the raw audio, and only wakes the full model when something interesting happens. If you run inference on every 20 millisecond frame around the clock, even a tiny model will drain a coin cell in days. If you gate inference behind a cheap trigger, you can stretch that to months.

Where the Cycles Actually Go

People usually assume the neural network dominates the power budget. On a Cortex-M4 running at a few tens of megahertz, that is often false. The feature extraction step, usually a mel-filterbank or MFCC computation, can cost as much as the model itself if you implement it naively. Floating point logarithms and trigonometric functions are expensive. Precompute your mel filterbank weights as constants, use a fast fixed-point log approximation, and avoid recomputing anything that does not change between frames. Small changes here routinely cut total energy per inference in half.

Memory access is the other hidden cost. Pulling weights from flash on every inference burns energy, so many runtimes keep the model in RAM if it fits. A keyword spotting model of a few tens of kilobytes can live comfortably in SRAM on many parts, which both speeds things up and reduces power. If your model is too big for RAM, consider whether you can stream weights or use a smaller architecture. On these devices, fitting in RAM is often a bigger win than shaving a few thousand MACs.

Then there is the microphone itself. A digital MEMS microphone with a proper low-power mode can be a significant fraction of your average current draw. Check the datasheet numbers at your actual sample rate. Sometimes choosing a lower sample rate, say 8 kHz instead of 16 kHz, is enough to matter, and keyword spotting models frequently work fine at the lower rate once you retrain them.

Measuring Instead of Guessing

The only honest way to tune a battery-powered design is to measure. Put a shunt resistor in series with your power supply and watch the current waveform on a scope while the device runs through its listen and infer cycle. You will see spikes you did not expect and idle currents you assumed were zero. Average current is what determines battery life, and it is dominated by the time spent in each state, not just the peak.

Build a simple state machine, log how many milliseconds you spend per second in microphone capture, feature extraction, inference, and sleep, and multiply by the measured current for each state. That spreadsheet will tell you more than any benchmark chart. Once you have it, you can make informed tradeoffs: a slightly larger model that runs less often may beat a smaller one that runs constantly.

None of this is glamorous, but it is the difference between a demo and a device you can actually ship on a battery.

Author

22 feb,2025

This incident opened my eyes to the value of Insurance in general, so I decided to examine my personal and business insurance.

Comment (3)

The payments are made directly from one person to another without passing through a central bank or clearing house.

22 feb,2025
Reply

A Practical Comparison of Lightweight Inference Runtimes for Android

23 feb,2025
Reply

Why Your Next Model Should Fit in a Megabyte: The Case for Tiny Neural Networks

23 feb,2025
Reply

Leave a Reply

Connect with us