
Book 32 of 50 · Free
TinyML: AI on Microcontrollers
3,656 words · 17 chapters · illustrated

Book 32 of 50 · Free
3,656 words · 17 chapters · illustrated
Book 32 of 50 — AstolixGen Learning Series For researcher and publication students

TinyML puts machine learning on microcontrollers with kilobytes of memory and milliwatts of power — no cloud, no GPU, no operating system. This book takes you from the "why" of TinyML through model compression, on-device inference frameworks, benchmarking, and deploying real models on boards like the ESP32. You will learn to design TinyML experiments you can publish: measuring the accuracy–energy–latency trade-off that is the heart of every TinyML paper.
Learning objectives: - Explain why TinyML matters: latency, privacy, energy, cost - Describe the constraints of MCU-class hardware (memory, compute, power) - Apply quantization, pruning, and knowledge distillation - Use TensorFlow Lite Micro and related frameworks - Benchmark models with MLPerf Tiny methodology - Deploy and measure a model on a real microcontroller - Design a publishable TinyML experiment
Cloud AI assumes connectivity, power, and money. TinyML assumes none of them. A microcontroller running a keyword-spotting model on a coin cell can listen for years; the same task via cloud needs a Wi-Fi radio, a data plan, and a server — and fails when the network does. Four forces drive TinyML. Latency: on-device inference answers in milliseconds with no round trip — critical for fall detection or industrial safety shutoffs. Privacy: audio and video never leave the device, sidestepping data-protection regulation. Energy: a radio transmission costs orders of magnitude more energy than local computation; TinyML keeps the radio off. Cost and scale: a $3 MCU can be deployed by the million where a $30 Linux board cannot.
The trade-off is capability: you get kilobytes of RAM, megahertz clocks, and no floating-point unit on the smallest parts. TinyML is therefore not "small deep learning" — it is a distinct discipline of co-designing models, software, and hardware under brutal constraints. For researchers, this constraint is the opportunity: every improvement in accuracy-per-milliwatt is a publishable result.
Example: A wildlife camera that streams video to the cloud needs cellular data and dies in days. A TinyML version runs person/animal detection on-device and only transmits a 100-byte "elephant detected" alert — lasting 6 months on the same battery and working where there is no coverage.
For your research: Frame TinyML papers around a constraint triplet: latency budget, energy budget, memory budget. State all three numerically in your abstract — reviewers check.
Key takeaway: TinyML trades model size for latency, privacy, energy, and cost wins — constraints are the research opportunity.
The typical TinyML target: an ARM Cortex-M4/M7 or ESP32-class chip with 256 KB–1 MB flash, 64–512 KB SRAM, clocks of 48–240 MHz, often no FPU, powered by battery. Compare with a phone (GBs of RAM, GPU) — you have roughly 1,000× less memory. Everything in TinyML flows from this.
Three resources bind you. Flash stores the model weights and code — your model must fit. SRAM holds activations (intermediate tensors) at runtime — often the tighter bound, because activations for a 96×96 image through a CNN can exceed the weights. Compute/energy bounds latency and battery: each inference costs multiply-accumulates (MACs), and energy per inference decides battery life.
Know your board's numbers before designing anything: flash size, SRAM size, clock, presence of FPU/DSP extensions (ARM CMSIS-NN/Helium accelerate quantized kernels dramatically), and sleep current. A paper that reports "runs on Cortex-M4 at 64 MHz" without SRAM/flash figures is incomplete.
Example: ESP32: 520 KB SRAM, 4 MB flash, 240 MHz dual-core, Wi-Fi/BT. Arduino Nano 33 BLE Sense: nRF52840, 256 KB SRAM, 1 MB flash, 64 MHz. The same keyword-spotting model fits easily on ESP32 but needs compression for the Nano — a comparison paper writes itself.
For your research: Always report the exact MCU, clock, memory map (model size, arena size), and whether DSP extensions were used. These details determine reproducibility.
Key takeaway: Flash, SRAM, and energy are the three walls — measure and report all three for your target MCU.
Quantization converts 32-bit floats to 8-bit (or smaller) integers. It cuts model size 4×, and on MCUs without FPUs it also speeds inference, because integer MACs are cheaper than emulated float. Two approaches: post-training quantization (PTQ) — quantize a trained float model, fast, small accuracy drop; and quantization-aware training (QAT) — simulate quantization during training, slower, usually recovers the accuracy PTQ loses.
The practical recipe: train in float, apply PTQ, measure accuracy on your test set. If the drop is acceptable (<1–2%), ship it. If not, use QAT. For MCUs, full-integer quantization (weights and activations int8, no float fallback) is required — "dynamic range" quantization that keeps some float ops will not run on integer-only kernels.
Watch for: operators your runtime doesn't support in int8 (check the op list), and activation ranges — a few outlier activations can destroy int8 resolution. Representative datasets for calibration matter; calibrate on data resembling deployment.
Example: A 2.1 MB float32 MobileNetV1 (0.25) for person detection → 550 KB int8 via PTQ, accuracy 80.2% → 79.6%. Fits ESP32 comfortably; with QAT recovers to 80.1%.
For your research: Quantization ablation studies (float32 vs int8 PTQ vs int8 QAT across 3+ models) are standard, expected, and citable components of TinyML papers. Report accuracy, size, latency, and energy for each.
Key takeaway: Int8 quantization is the single highest-leverage TinyML technique — 4× smaller, often faster, tiny accuracy cost.
Pruning removes weights (or whole channels/filters) that contribute little. Unstructured pruning (individual weights) compresses well but needs sparse kernels to speed up; structured pruning (removing channels) keeps dense math and directly cuts latency on MCUs. Iterative prune → fine-tune beats one-shot pruning.
Knowledge distillation trains a small "student" model to mimic a large "teacher" (matching soft outputs, not just hard labels). The student learns the teacher's generalization, often beating a small model trained from scratch by 2–5%. Distillation is almost free accuracy for TinyML — train big on the server, deploy small on the device.
Neural Architecture Search (NAS) automates finding MCU-friendly architectures. MCUNet (Lin et al., NeurIPS 2020) used TinyNAS + TinyEngine to co-design models for 256 KB SRAM, achieving record ImageNet accuracy on microcontrollers. You don't need to invent NAS — use published tiny architectures (MobileNetV1 0.25, MicroNets, MCUNet) as starting points and cite them.
Example pipeline: Teacher ResNet50 (95% on your task) → distill to MobileNetV1-0.25 student (91%) → structured prune 30% of channels + fine-tune (90.2%) → int8 QAT (89.8%, 480 KB). Each step measured and reported.
For your research: Compression pipelines make excellent "methods" sections: each stage is ablated, each trade-off plotted. Compare your pipeline against at least one published tiny model as baseline.
Key takeaway: Prune structure not just weights, distill from big teachers, and start from published tiny architectures.
TensorFlow Lite Micro (TFLM) (David et al., MLSys 2021) is the standard inference runtime: an interpreter with a tiny binary (~20 KB), no dynamic allocation (all memory from a fixed "arena" you size), and kernels optimized for Cortex-M (CMSIS-NN). The workflow: train in TensorFlow/PyTorch → convert to TFLite (with quantization) → embed the .tflite as a C array → run the TFLM interpreter with a sized tensor arena.
Key practical points: the arena size must cover the largest activation tensor plus overhead — too small and inference fails at runtime; profile with the recording allocator. Only include the operators you use (OpResolver) to keep binary small. CMSIS-NN integration gives 2–4× speedup on Cortex-M — always enable it when available and report whether it was on.
Alternatives: STM32Cube.AI (converts to optimized C for STM32), Edge Impulse (end-to-end studio: data → training → deployment, excellent for fast prototyping and teaching), TinyEngine (from the MCUNet work, memory-scheduled inference). For research papers, TFLM is the most cited and most reproducible choice.
Example: Person-detection model: 312 KB int8 TFLite → arena 180 KB → total SRAM 492 KB + overhead fits ESP32's 520 KB with room for application code. Binary with 12 ops + CMSIS-NN: 68 KB.
For your research: Report converter version, quantization scheme, arena size, op list, and CMSIS-NN status. These are the "implementation details" reviewers need to reproduce.
Key takeaway: TFLM + int8 + CMSIS-NN is the default stack — size the arena, trim the ops, report everything.
You cannot claim "efficient" without a benchmark. MLPerf Tiny (Banbury et al., NeurIPS 2021) is the community standard: four tasks (keyword spotting on Speech Commands, visual wake words on COCO-derived data, image classification on CIFAR-10, anomaly detection on ToyADMOS), with strict rules on training data and measurement of latency and energy. Running your model on MLPerf Tiny tasks makes your numbers comparable with the literature — this is the single most impactful thing you can do for a TinyML paper's credibility.
Beyond MLPerf: Speech Commands (keyword spotting), Visual Wake Words (person detection), CIFAR-10 (sanity-check classification). For applied papers, collect your own dataset but also evaluate on a public one — reviewers trust public benchmarks and question private-only results.
Example: "Our model reaches 91.2% on Visual Wake Words (vs 88.4% MicroNet baseline) at 38 ms and 4.1 mJ per inference on Cortex-M7 @ 216 MHz" — every number comparable, every claim checkable.
For your research: If you can't run full MLPerf Tiny, run at least one of its tasks. State deviations from the rules explicitly. Never report accuracy without latency and energy on named hardware.
Key takeaway: MLPerf Tiny is the credibility standard — benchmark on it, report latency + energy + accuracy together.
TinyML evaluation is three-dimensional. Latency: time per inference, measured on-device with cycle counters or GPIO toggling (not on your laptop). Report mean and worst-case; real-time tasks need deadlines, not averages. Energy: joules per inference, measured with a power profiler (Nordic PPK2, Joulescope) or a shunt + INA219. Report at the system level (including sensor and radio if they run during inference). Memory: flash (model + code) and peak SRAM (arena) — report both, since either can block deployment.
Methodology pitfalls: measuring on a development board with the debugger attached (changes power), reporting laptop inference time, single-run numbers without variance, and forgetting idle/sleep power in duty-cycled lifetime estimates. Battery-life claims must include the full duty cycle: sense → infer → (maybe) transmit → sleep.
Example lifetime math: Inference 4.1 mJ every 10 s + sleep 15 µA → average power ≈ 0.45 mW → on a 1000 mAh coin cell ≈ 2.5 years. Show this math in applied papers; reviewers love it and it proves real-world viability.
For your research: A "measurement methodology" subsection with instrument names, sample counts, and variance is expected in top TinyML venues. Publish raw measurement logs.
Key takeaway: Measure latency, energy, and memory on-device with proper instruments — and show the battery-life math.
The ESP32 is the most accessible TinyML target (cheap, Wi-Fi, large community). Deployment path: (1) Train/quantize model → .tflite. (2) Convert to C array (xxd -i model.tflite). (3) In ESP-IDF or Arduino: allocate arena (e.g., 200 KB in PSRAM or SRAM), register ops, load model, feed preprocessed sensor data, invoke, read outputs. (4) Add application logic: duty cycling (deep sleep between inferences), and only power the radio when an event is detected.
Practical ESP32 notes: use PSRAM for large arenas but note its higher latency; keep the Wi-Fi radio off during inference for clean energy measurements; use the ULP coprocessor or timer wakeup for duty cycling; partition flash so OTA updates don't brick the model slot.
Example: Keyword spotting on ESP32: 1 s audio window → MFCC features (on-device, ~12 ms) → int8 model (9 ms) → total 21 ms, 520 µJ per check, checking every 2 s → average under 1 mW. Detected keyword triggers Wi-Fi upload of a 5 s clip.
For your research: ESP32 deployment papers should report: framework version, partition table, arena placement (SRAM vs PSRAM), radio state during measurement, and deep-sleep current. These details separate reproducible work from demos.
Key takeaway: ESP32 + TFLM + duty cycling is the canonical deployment — document every configuration choice.
Most TinyML applications use one of three modalities. Audio (keyword spotting, event detection): features are MFCCs or log-mel spectrograms computed on-device; models are small CNNs or DS-CNNs; the challenge is always-on power (duty-cycle the feature extractor). Vision (person detection, gesture): input is heavily downsampled (96×96 grayscale common); models are MobileNet-class; the challenge is SRAM for activations — design networks with small early feature maps. IMU (activity recognition, anomaly detection): 1D time series; tiny 1D-CNNs or autoencoders; the challenge is labeling real-world motion data.
Choose the modality that matches your problem's physics: a microphone hears through walls (privacy issue), a camera needs light, an IMU needs contact. Multimodal fusion (audio + IMU) is an under-explored research area with easy wins.
Example: A fall detector using only IMU: 3-axis accel @ 50 Hz → 2 s windows → 1D-CNN autoencoder (38 KB) → anomaly score. No camera (privacy), no microphone, 6-month battery. Published results focus on false-alarm rate, the metric that matters for adoption.
For your research: Pick one modality and own it. Cross-modality comparisons on the same task (audio vs IMU for machine anomaly detection) are clean, useful papers.
Key takeaway: Audio, vision, IMU — match modality to physics; report the metric users care about (false alarms, not just accuracy).
Applied TinyML papers win by solving a real problem under real constraints. Agriculture: pest detection from leaf images on a solar node; irrigation triggers from soil + weather models. Health/wearables: arrhythmia detection on ECG, cough detection, fall detection — privacy-sensitive, so on-device is a feature. Environment: air-quality inference from low-cost sensors, wildlife monitoring, flood sensing.
The applied-paper formula: (1) real problem with a stakeholder, (2) why cloud fails here (no connectivity/power/budget), (3) dataset collected in the field (not just lab), (4) model meeting the triple budget, (5) field validation with honest failure analysis. Step 5 is what separates publications from demos — report what broke.
Example: A tea-estate pest detector: 4,000 field leaf photos (not PlantVillage lab images), int8 MobileNetV1-0.25 at 84% field accuracy, solar + ESP32, 3-month deployment, failure analysis showing accuracy drop in rain (water droplets) — the honest failure section became the paper's most-cited part.
For your research: Field data + failure analysis is the applied TinyML paper. Lab-only results on public datasets are fine for methods papers but weak for applied ones.
Key takeaway: Real problem, real constraints, field data, honest failures — the applied TinyML formula.
Open problems, roughly ordered by accessibility: (1) Better benchmarks — MLPerf Tiny covers 4 tasks; new tasks (tiny object detection, on-device learning) need benchmarks. (2) NAS for new hardware — most NAS targets Cortex-M; RISC-V and NPUs need their own. (3) On-device training/federated TinyML — personalization without the cloud is barely explored. (4) Toolchain gaps — debugging, profiling, and CI for TinyML are primitive; systems papers welcome. (5) Security — model extraction and adversarial examples on MCUs are under-studied. (6) Multimodal tiny fusion — principled fusion under KB budgets.
Read the venues: MLSys, NeurIPS (Datasets/Benchmarks, Efficient ML workshops), IEEE Micro, ACM SenSys/IMWUT for applied work, TinyML Research Symposium. Follow the MCUNet/MicroNet line and the MLPerf Tiny leaderboard — that is the state of the art you must beat or build on.
Example contribution map: "We extend MLPerf Tiny with a fifth task (tiny object detection on 96×96), provide baselines, and show current methods lose 12% mAP under the 256 KB SRAM cap" — a benchmark paper, highly citable.
For your research: Pick one open problem, reproduce the closest prior work first (this validates your setup), then improve one axis. Reproduction + one improvement = a solid first TinyML paper.
Key takeaway: Benchmarks, NAS, on-device learning, tooling, security — pick one gap, reproduce first, then improve.
Assemble the pipeline:
Common rejection reasons: no baseline comparison, laptop-only measurements, missing memory numbers, private dataset with no public-task validation, and claiming "real-time" without a deadline analysis.
For your research: This chapter mirrors your assignment structure: problem + lit review (Assignment 1), method (Assignment 2), results (Assignment 3), presentation (Assignment 4). A TinyML study fits it perfectly.
Key takeaway: Triple budget + reproduced baseline + one improvement + full measurement + artifact = a publishable TinyML paper.
| # | Chapter | Core idea | Research use |
|---|---|---|---|
| 1 | Why TinyML | Latency/privacy/energy/cost | State constraint triplet numerically |
| 2 | MCU hardware | Flash/SRAM/compute walls | Report exact MCU + memory map |
| 3 | Quantization | int8 PTQ/QAT | Ablation: float32 vs PTQ vs QAT |
| 4 | Pruning/distill/NAS | Structured pruning, KD, tiny archs | Pipeline ablations vs published baselines |
| 5 | TFLM stack | Arena, ops, CMSIS-NN | Report converter/arena/op versions |
| 6 | MLPerf Tiny | 4 standard tasks | Benchmark for credibility |
| 7 | Measurement | Latency/energy/memory on-device | Instruments, variance, battery math |
| 8 | ESP32 deploy | Duty cycling, radio discipline | Full config disclosure |
| 9 | 3 senses | Audio/vision/IMU trade-offs | Modality comparison papers |
| 10 | Applied domains | Field data + failure analysis | Real deployments, honest failures |
| 11 | Open problems | Benchmarks, NAS, security gaps | Contribution mapping |
| 12 | Study design | Baseline + one improvement + artifact | Assignment-to-paper pipeline |
[1] C. R. Banbury et al., "MLPerf Tiny Benchmark," in Proc. 35th Conf. Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2021. [2] R. David et al., "TensorFlow Lite Micro: Embedded Machine Learning for TinyML Systems," in Proc. 4th Conf. Machine Learning and Systems (MLSys), 2021. [3] J. Lin et al., "MCUNet: Tiny Deep Learning on IoT Devices," in Proc. 34th Conf. Neural Information Processing Systems (NeurIPS), 2020. [4] J. Lin et al., "MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning," in Proc. 35th Conf. Neural Information Processing Systems (NeurIPS), 2021. [5] A. Gholami et al., "A Survey of Quantization Methods for Efficient Neural Network Inference," arXiv:2103.13630, 2021. (survey) [6] W. Shi et al., "Edge Computing: Vision and Challenges," IEEE Internet of Things Journal, vol. 3, no. 5, pp. 637–646, 2016. [7] L. Atzori, A. Iera, and G. Morabito, "The Internet of Things: A survey," Computer Networks, vol. 54, no. 15, pp. 2787–2805, 2010. [8] P. Warden and D. Situnayake, TinyML: Machine Learning with TensorFlow Lite on Arduino and Ultra-Low-Power Microcontrollers, O'Reilly Media, 2019. (book) [9] S. S. Choudhary et al., "A survey on TinyML," IEEE Access, 2022. (verify current citation before use) [10] M. Satyanarayanan, "The Emergence of Edge Computing," Computer, vol. 50, no. 1, pp. 30–39, 2017.
End of Book 32. Next: Book 33 — Edge AI vs Cloud AI.