Contents
Quick, necessary correction before anything else: the original Jetson Nano stopped being manufactured back in 2022. If you already own one, it still works and NVIDIA is committed to software support through 2027 — genuinely fine to keep learning on. But if you're about to buy one specifically for a new project in 2026, you'd actually be buying its successor, the Jetson Orin Nano, which delivers up to 80x the AI performance in the same compact form factor. NVIDIA just pushed this line forward again on August 25, 2026 with the Orin Nano 2, doubling inference throughput over last year's model.
This tutorial covers both honestly: the setup process that still applies if you have an original Nano, and the reasons you'd want the current Orin Nano generation instead if you're starting fresh. This is genuinely the platform this whole edge AI series has been building toward — the top of the hardware spectrum from the earlier hardware guide, capable of running things no microcontroller or bare Raspberry Pi could touch.
By the end of this guide, you'll understand which Jetson generation actually fits your situation, have one flashed and running, and have your first deep learning inference project working on real GPU-accelerated edge hardware. IMO, the jump from Raspberry Pi's CPU-bound inference to a Jetson's actual CUDA cores is a genuinely different tier of capability, not just an incremental upgrade :)
Which Jetson Should You Actually Buy in 2026?
This matters enormously before you spend any money, since the naming across generations genuinely confuses people.
- Original Jetson Nano — production ended in 2022, delivers 472 GFLOPS, software supported until 2027. Only worth using if you already own one. It's a legitimate, cheap way to learn AI fundamentals, but don't buy one new for a serious 2026 project.
- Jetson Orin Nano Super (introduced late 2024, still widely available) — 67 sparse INT8 TOPS, 8GB LPDDR5, priced around $249. This is genuinely the current mainstream recommendation for most people starting a new project today.
- Jetson Orin Nano 2 (announced August 25, 2026) — 78 TOPS, an 8-core Arm CPU, doubling inference performance over the Orin Nano Super while cutting power draw 40% at matched performance. The newest option, worth checking availability and pricing if you're buying right now.
The practical guidance: if you already have an original Nano, this tutorial's setup process still applies — use it for learning, not for anything demanding real throughput. If you're buying new, get an Orin Nano-generation board; the original Nano genuinely can't run the vision-language and generative workloads that make this hardware tier interesting in 2026.
What Makes Jetson Different From Everything Else in This Series
Recall the Raspberry Pi tutorial's core limitation: bare CPU inference struggled with real-time vision, requiring a bolt-on accelerator (Coral, Hailo) to hit usable frame rates. Jetson boards skip that problem entirely by design.
- A genuine CUDA-capable GPU sits on the same board — 128 CUDA cores on the original Nano, 1,024 Ampere-architecture cores plus 32 Tensor Cores on the Orin Nano Super — not a bolt-on accelerator, but a first-class part of the SoC.
- JetPack SDK bundles CUDA, cuDNN, and TensorRT together — the exact same NVIDIA acceleration stack referenced in the ONNX Runtime article's CUDAExecutionProvider and TensorRTExecutionProvider sections, running natively on this hardware rather than through translation layers.
- This is the platform the edge AI hardware guide specifically named as the tier for generative and transformer-based edge workloads — vision-language models, small LLMs, diffusion models — categorically beyond what Raspberry Pi or any microcontroller can realistically handle.
Setting Up: Flashing JetPack
The setup process genuinely differs based on which generation you have, so confirm your board before starting.
For the Original Jetson Nano
# Download JetPack 4.6.1 SD card image from NVIDIA's Jetson Nano page
# Flash to a Class 10, 32GB+ microSD card using Balena Etcher
Insert the flashed card, connect an HDMI monitor, USB keyboard/mouse, and a 5V/4A power supply, then boot through the initial Ubuntu 18.04-based setup. The whole process genuinely takes under 30 minutes for anyone who's flashed an SD card before.
For Jetson Orin Nano (Super or 2)
# Download and run the NVIDIA SDK Manager on an Ubuntu 20.04 or 22.04 host machine
# Force the board into recovery mode, then flash via SDK Manager
One genuinely important gotcha worth flagging upfront: flashing an Orin-generation board typically requires your host machine to run a specific Ubuntu version — mismatches here are a common source of confusing flash failures. Check NVIDIA's current JetPack release notes for your exact board before starting, since supported host OS versions do shift between JetPack releases.
Once flashed, confirm your environment:
sudo apt-get update && sudo apt-get upgrade -y
sudo apt-get install python3 python3-dev python3-pip python3-venv -y
Your First Deep Learning Project: Hello AI World
NVIDIA's own jetson-inference repository (maintained by dusty-nv) remains the standard "hello world" entry point for Jetson boards across every generation — genuinely the same recommended starting project whether you're on an original Nano or the newest Orin Nano 2.
git clone --recursive https://github.com/dusty-nv/jetson-inference
cd jetson-inference
mkdir build
cd build
cmake ../
make -j$(nproc)
sudo make install
This builds a collection of ready-to-run inference examples — image classification, object detection, and pose estimation (PoseNet) among them — genuinely useful for confirming your whole CUDA/TensorRT stack works correctly before writing any of your own code.
cd build/aarch64/bin
./imagenet images/orange_0.jpg output.jpg
Watching this correctly classify an image on your first run is the equivalent moment to the Raspberry Pi tutorial's first classification result — except now you're seeing genuine GPU-accelerated inference, not CPU inference straining to keep up.
TensorRT: Where Jetson's Real Advantage Shows Up
Recall from the ONNX Runtime article: TensorRT is NVIDIA's specific graph-optimization execution provider, squeezing extra performance from NVIDIA hardware beyond generic CUDA acceleration. On Jetson, this isn't an optional add-on — it's the standard path to genuinely usable performance.
import tensorrt as trt
logger = trt.Logger(trt.Logger.WARNING)
builder = trt.Builder(logger)
network = builder.create_network()
parser = trt.OnnxParser(network, logger)
with open("model.onnx", "rb") as f:
parser.parse(f.read())
config = builder.create_builder_config()
config.set_flag(trt.BuilderFlag.FP16) # or INT8 for further speedup
engine = builder.build_engine(network, config)
Notice BuilderFlag.FP16 — this is exactly the quantization concept from earlier in this series, applied through TensorRT's own optimization path rather than GGUF or TFLite's. Converting a trained ONNX model into an optimized TensorRT engine typically delivers a genuinely substantial speedup over running the same model through generic inference, precisely because TensorRT restructures the computation graph specifically for your board's exact GPU architecture.
Real-Time Object Detection: Applying the YOLO Article Here
Recall the YOLO26-on-Raspberry-Pi article's realistic 15-20 FPS ceiling on bare Pi CPU. Jetson hardware changes that equation substantially.
- On an Orin Nano Super, YOLO-class models run considerably faster than Pi's CPU-bound NCNN export — genuinely enabling smoother real-time tracking, higher resolution input, or multiple simultaneous camera streams that would overwhelm a Pi.
- YOLO26 explicitly lists Jetson as a target export platform, alongside the Raspberry Pi and mobile NPU targets covered in that earlier article — the same ultralytics package and export pipeline applies, just targeting TensorRT as the export format instead of NCNN.
from ultralytics import YOLO
model = YOLO("yolo26n.pt")
model.export(format="engine") # TensorRT export, Jetson-specific
trt_model = YOLO("yolo26n.engine")
This is genuinely the same code shape from the Raspberry Pi YOLO article, with format="engine" replacing format="ncnn" — the export target changes to match the hardware, but the workflow you already learned transfers directly.
Running Generative AI Locally: The Jetson AI Lab
This is genuinely the capability that separates Jetson from every other device covered in this series' edge AI arc. The Jetson Generative AI Lab provides containerized installs specifically for running LLMs, VLMs, and Whisper-class models on this hardware, with real, current documentation rather than a bare board you're left to configure alone.
- NVIDIA's own materials note that "small and medium frontier models have reached the accuracy of last year's largest frontier models" — meaning genuinely capable models now fit within an Orin Nano's memory and compute budget in a way that wasn't true even a year or two ago.
- The Orin Nano 2 specifically runs open models like NVIDIA Nemotron, Gemma, and Qwen optimized for memory-efficient edge inference — connecting directly back to the local LLM tooling covered earlier in this series, just running on dedicated edge hardware instead of a laptop.
- This is genuinely where the local LLM series and the edge AI series converge — everything you learned about GGUF, quantization, and llama.cpp applies conceptually here too, with Jetson's TensorRT-LLM providing the platform-specific optimization layer.
If you're running generative workloads on Jetson, the RTX 5070 and RTX 5080 serve as the desktop counterparts — the same TensorRT optimization stack applies, just with more VRAM and raw throughput for larger model variants.
Realistic Performance Expectations by Use Case
Worth setting concrete expectations rather than assuming marketing TOPS numbers translate directly to your specific project.
- YOLO-class object detection: genuinely smooth real-time performance on Orin-generation hardware, single or multiple camera streams depending on model size and board tier.
- Small local LLMs and VLMs: increasingly practical on Orin Nano-class hardware specifically because "small" models have gotten meaningfully more capable, not because the hardware got dramatically bigger.
- Local Stable Diffusion: SD 1.5 works reasonably (roughly 30 seconds per 512x512 image on entry-level Orin hardware); SDXL genuinely struggles against an 8GB memory ceiling — fine for prototyping and thumbnails, not production image generation.
- Original Jetson Nano, for comparison: handles roughly 90% of purely educational workloads adequately, but shouldn't be expected to run anything in the generative AI category this article covers for Orin-tier hardware.
Want to Go Deeper?
If the TensorRT optimization or CUDA kernel concepts clicked and you want to dig into the theory behind GPU-accelerated inference and edge deployment, Educative's Machine Learning path covers hardware-specific optimization in detail alongside the broader deep learning deployment landscape — worth exploring if you're building production Jetson pipelines.
Common Mistakes People Make
- Buying an original Jetson Nano new in 2026 for anything beyond pure learning. It's discontinued and genuinely can't handle current generative workloads — the Orin Nano generation is the correct purchase for an active project.
- Mismatching your flashing host machine's Ubuntu version. This is a genuinely common, confusing failure point — check your specific JetPack version's documented host OS requirement before starting.
- Skipping TensorRT conversion and running raw ONNX or PyTorch models directly. You're leaving real performance on the table — the graph optimization TensorRT provides is a substantial part of why this hardware tier is worth its cost over a bare Raspberry Pi.
- Expecting SDXL-quality image generation on entry-level Orin hardware. The 8GB memory ceiling genuinely limits this specific use case — SD 1.5 is the realistic target, not SDXL.
- Assuming TOPS figures directly compare across model families. Sparse INT8 TOPS, dense TOPS, and different precision modes aren't directly comparable numbers — check what specific measurement a spec sheet is actually quoting before comparing boards.
Where This Fits With the Rest of This Series
This genuinely closes out the edge AI hardware spectrum from earlier in this series — Arduino handled kilobyte-scale TinyML, Raspberry Pi handled genuine Python and moderate vision workloads, and Jetson is the tier where transformer-based and generative AI actually becomes practical at the edge. The quantization, pruning, and distillation techniques from the model compression articles all apply here too, now paired with TensorRT's hardware-specific optimization layer rather than GGUF or TFLite alone.
Wrapping This Up
Getting started with Jetson in 2026 genuinely means picking the Orin Nano generation over the original Nano for any new project — the naming is confusing, but the capability gap between them is substantial, not incremental. JetPack's bundled CUDA/cuDNN/TensorRT stack, combined with the Jetson AI Lab's containerized generative AI tooling, makes this the only platform in this entire series capable of running genuinely current LLM and vision-language workloads locally on edge hardware.
Remember that an original Jetson Nano is a fine learning tool if you already own one, but a poor purchase decision for a new 2026 project, and that TensorRT conversion is genuinely not optional if you want this hardware tier's real performance advantage over Raspberry Pi. FYI, the fact that NVIDIA specifically highlighted small models now matching last year's largest frontier models' accuracy is a real signal that this hardware tier's practical capability is still expanding month to month, not a settled endpoint :)
Now go run the jetson-inference PoseNet example on whatever board you've got and watch real-time pose estimation happening entirely on-device. That's genuinely the moment this hardware tier's GPU-accelerated advantage over everything else in this series becomes viscerally obvious.