Sam Austin AI

TinyML on Arduino: Your First Machine Learning Model on a Microcontroller (2026)

September 7, 2026 12 min read Sam Austin
Contents
TinyML on Arduino First Machine Learning Model Microcontroller Tutorial
TinyML on Arduino First Machine Learning Model Microcontroller Tutorial

Figure 1: Running neural networks on Arduino microcontrollers — machine learning at the edge

Here's a fact that genuinely reframes what "small" means in machine learning: Google's wake-word detector — the model listening for "Hey Google" — is roughly 14 kilobytes. Not 14 megabytes. Fourteen kilobytes, small enough to run on a chip that costs less than a sandwich. Ever wondered how your phone hears a wake word without streaming raw audio to a server constantly? This is exactly the field responsible.

Building on the edge AI overview and hardware guide from earlier in this series, this is the actual hands-on project — training a genuinely tiny neural network, converting it, and getting it running on real Arduino hardware. No cloud, no GPU, just a microcontroller doing inference on its own.

By the end of this tutorial, you'll have trained a small model in Python, converted it to run on-device, and deployed it to an Arduino, watching a physical output respond to a neural network's prediction in real time. IMO, there's something genuinely delightful about a "hello world" project that ends with a blinking LED controlled by machine learning rather than a print statement :)

The Classic TinyML "Hello World": Predicting a Sine Wave

Every TinyML tutorial converges on essentially the same starting project, and for good reason: train a tiny neural network to approximate the sine function, then use its output to control something physical — typically an LED's brightness, pulsing smoothly like a heartbeat.

This is, admittedly, one of the most inefficient, roundabout ways to calculate a sine wave you'll ever see — Arduino has a built-in sin() function that does this instantly with zero training required. That's exactly the point. The task is simple enough that you can focus entirely on the actual TinyML pipeline — train, convert, deploy — without a complicated model architecture getting in the way.

Step 1: Training the Model in Python

We'll build a small, fully connected neural network — three layers — trained as a regression model to predict sine wave output from an input angle.

import tensorflow as tf
import numpy as np

x_values = np.random.uniform(low=0, high=2*np.pi, size=1000).astype(np.float32)
np.random.shuffle(x_values)
y_values = np.sin(x_values)

model = tf.keras.Sequential([
    tf.keras.layers.Dense(16, activation='relu', input_shape=(1,)),
    tf.keras.layers.Dense(16, activation='relu'),
    tf.keras.layers.Dense(1)
])

model.compile(optimizer='rmsprop', loss='mse', metrics=['mae'])
model.fit(x_values, y_values, epochs=500, batch_size=64, validation_split=0.2)

Notice how small this network genuinely is — two hidden layers of just 16 neurons each. This isn't a scaled-down version of a "real" model; it's already the right size for the task, which is exactly the mindset TinyML demands: don't build bigger than the problem needs, because on a microcontroller, every unnecessary parameter has a real memory cost.

Step 2: Converting to TensorFlow Lite

A trained Keras model can't run on a microcontroller directly — it needs conversion into a format designed specifically for resource-constrained inference.

converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
tflite_model = converter.convert()

with open("sine_model.tflite", "wb") as f:
    f.write(tflite_model)

The TensorFlow Lite model is stored as a FlatBuffer — a format specifically useful for reading large chunks of data incrementally, rather than needing to load everything into RAM at once. That Optimize.DEFAULT flag triggers quantization automatically, shrinking the model further using exactly the same underlying principle from the edge AI hardware article — trading a small amount of precision for a meaningfully smaller footprint.

Step 3: Converting the Model Into C Code

Here's the genuinely unusual step that has no equivalent in normal Python ML work: your microcontroller doesn't have a filesystem the way your laptop does, so the model needs to become part of the actual program, embedded directly as a C array.

xxd -i sine_model.tflite > sine_model.h

This produces a header file containing your entire model as a byte array — something like:

unsigned char sine_model_tflite[] = {
  0x1c, 0x00, 0x00, 0x00, 0x54, 0x46, 0x4c, 0x33, ...
};
unsigned int sine_model_tflite_len = 2656;

This is genuinely the moment TinyML feels different from any other ML workflow you've touched — your trained neural network is now, quite literally, hardcoded into your firmware, compiled directly into the program that runs on the chip.

Step 4: Setting Up the Arduino Side

You'll need the Arduino_TensorFlowLite library (TensorFlow Lite for Microcontrollers) installed through the Arduino IDE's Library Manager, plus a supported board — the Arduino Nano 33 BLE Sense is the standard choice, combining a capable Cortex-M4 microcontroller with onboard sensors in one package.

#include <TensorFlowLite.h>
#include "sine_model.h"
#include "tensorflow/lite/micro/all_ops_resolver.h"
#include "tensorflow/lite/micro/micro_interpreter.h"
#include "tensorflow/lite/schema/schema_generated.h"

constexpr int kTensorArenaSize = 2 * 1024;
uint8_t tensor_arena[kTensorArenaSize];

const tflite::Model* model = tflite::GetModel(sine_model_tflite);
tflite::AllOpsResolver resolver;
tflite::MicroInterpreter interpreter(model, resolver, tensor_arena, kTensorArenaSize);

That tensor_arena is worth understanding specifically — it's a fixed block of memory the interpreter uses for all its intermediate calculations during inference. Getting this size wrong is one of the most common TinyML stumbling blocks: too small, and inference fails; unnecessarily large, and you've wasted scarce RAM the rest of your program might need.

Running Inference and Controlling the LED

void loop() {
  float x = get_current_angle();  // your input value

  TfLiteTensor* input = interpreter.input(0);
  input->data.f[0] = x;
  interpreter.Invoke();

  TfLiteTensor* output = interpreter.output(0);
  float y_value = output->data.f[0];

  int brightness = (int)(255 * (y_value + 1) / 2);
  analogWrite(LED_PIN, brightness);
}

Watch that LED pulse smoothly, and you're looking at a neural network's prediction driving physical hardware in real time — genuinely the same core loop as every larger inference pipeline you've built in this series, just running entirely inside a chip smaller than your thumbnail.

Beyond Sine Waves: What This Actually Enables

The sine wave project is deliberately trivial, but the exact same five-stage pipeline — build, train, convert, deploy, run — scales up to genuinely useful applications once you swap in real sensor data.

Gesture recognition: feeding the Nano 33 BLE Sense's onboard IMU data through a trained classifier, detecting specific motions like a "magic wand" flick or a wave.

Wake-word detection: the same category of model responsible for "Hey Google," trained on audio samples through the onboard microphone.

Person detection: a lightweight vision model running on camera-equipped boards, classifying whether a person is present in frame.

Simple sensor-based classification: temperature, motion, or environmental anomaly detection — reading a sensor value, running it through a trained model, and reacting based on the prediction, directly on-device.

The architecture and training process for each of these follows the same shape you just built — different input data, different model complexity, but the same build-train-convert-deploy pipeline underneath.

If you want to take this further into robotics, the robotics kits guide covers Arduino and Raspberry Pi hardware for RL experimentation — adding TinyML sensor intelligence to those platforms is a genuinely achievable next step.

A Genuinely Common Pitfall: The Tensor Arena

Worth calling out specifically since it trips up nearly everyone on a first TinyML project: microcontrollers have severely limited RAM, and TensorFlow Lite requires a chunk of it specifically set aside for the "tensor arena" used during inference computation.

Underestimating this size is a genuinely common recurring mistake — your program compiles fine, but inference silently fails or the board crashes at runtime because the arena couldn't hold the intermediate calculations a given model actually needs.

The fix is usually iterative: start with a reasonable estimate, and increase it if inference fails, rather than trying to calculate the exact required size analytically upfront.

This is a genuinely different debugging category from anything in normal Python ML work — you're not debugging a training curve, you're debugging a memory allocation on hardware with kilobytes, not gigabytes, to spend.

Testing Under Real Conditions

Here's advice worth taking seriously once you move past the sine wave demo into a genuinely sensor-driven project: real-world performance often differs meaningfully from training metrics.

Test extensively under varied conditions — different hand positions and movement speeds for gesture recognition, different lighting for vision tasks, different background noise for audio.

If accuracy degrades under real conditions, the fix is usually collecting additional training data specifically from those problematic conditions, not just tweaking model architecture.

This mirrors the sim-to-real gap discussed earlier in this series for robotics — a model that looks great in your clean training data can genuinely behave differently once real-world noise and variation enter the picture.

Common Mistakes People Make

Underestimating the tensor arena size. This is the single most common first-project stumbling block — start with a working estimate and adjust rather than guessing perfectly upfront.

Building a model too large for the target microcontroller before checking memory constraints. Confirm your board's actual RAM and flash budget before training something ambitious.

Skipping quantization during conversion. The Optimize.DEFAULT flag genuinely matters — without it, you're deploying a needlessly large model for no real accuracy benefit at this scale.

Assuming training-set accuracy predicts real-world performance. Physical sensors introduce noise and variation training data often doesn't fully capture — test on real hardware under real conditions before trusting the numbers.

Not checking board compatibility before starting. TensorFlow Lite for Microcontrollers support varies by board — confirm your specific Arduino model is actually supported before investing time in a project it can't run.

Where This Fits With Everything Else in This Series

If you've worked through the edge AI overview and hardware comparison articles, this tutorial is genuinely the hands-on payoff those two were building toward. The Arduino Nano 33 BLE Sense recommended in the hardware article is the exact board this tutorial deploys to, and the quantization concept from the edge AI overview is the literal Optimize.DEFAULT line in your conversion step.

This is also worth connecting back to the robotics kits article — if you've got an Arduino-based robot from that piece, a TinyML classifier like this is a genuinely achievable way to give it actual on-device sensor intelligence, rather than relying entirely on a tethered laptop for every decision.

Next Steps: Build Smarter Embedded Systems

Ready to take your embedded ML skills further? Educative offers interactive courses on edge AI, TensorFlow Lite, and building production ML applications — practice in sandboxed environments and learn by building real projects.

Wrapping This Up

Getting your first model running on Arduino walks through the complete TinyML pipeline in miniature: train a genuinely small network in Python, convert it to TensorFlow Lite with quantization, embed it as a C array, and run inference directly on microcontroller hardware with a physical output responding in real time. The sine wave example is deliberately trivial, but the exact same five stages scale directly to gesture recognition, wake-word detection, and sensor-based classification.

Remember that tensor arena sizing is the most common first-project stumbling block, and that real-world testing under varied conditions matters more than training accuracy alone — physical sensors introduce noise your clean training data won't fully anticipate. FYI, the fact that a functioning neural network can run in roughly 2KB of working memory is genuinely one of the more mind-bending facts in this whole field once it actually clicks through building it yourself :)

Now go swap the sine wave for real accelerometer data from your board's onboard IMU and try training a basic gesture classifier next. That's the genuinely satisfying next step once the core pipeline here feels familiar rather than mysterious.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles