Contents
Figure 1: ONNX Runtime — deploy the same model file across NVIDIA GPUs, Intel CPUs, Apple Silicon, and more
Here's a problem this whole series has quietly danced around: you trained a model in PyTorch, but your deployment target expects TensorFlow. Or your team standardized on scikit-learn, but production needs C++. Every framework speaks its own language, and every hardware vendor has its own acceleration API. ONNX Runtime exists specifically to make that mismatch stop mattering.
This is genuinely the framework-agnostic counterpart to the LiteRT article — where LiteRT is Google's answer specifically for mobile and embedded, ONNX Runtime is Microsoft's answer for literally everything else: cloud servers, desktop apps, edge devices, and mobile too, all through one universal model format and one consistent API regardless of what actually executes it underneath.
By the end of this guide, you'll understand what ONNX actually is, how Execution Providers let the exact same model file run optimally on wildly different hardware, and the practical conversion and deployment pipeline. IMO, the "write once, run anywhere" pitch is one of those claims that's usually exaggerated in software — this is a genuinely rare case where it mostly holds up :)
What ONNX Runtime Actually Is
ONNX Runtime (ORT) is a cross-platform inference and training accelerator built around the ONNX format — a portable computation graph representation that models trained in wildly different frameworks can all convert into.
It's compatible with deep learning frameworks like PyTorch and TensorFlow/Keras, as well as classical machine learning libraries like scikit-learn, LightGBM, and XGBoost — genuinely broader framework coverage than the TensorFlow/PyTorch/JAX scope of LiteRT.
ONNX Runtime's inference APIs have been stable and production-ready since the 1.0 release in October 2019 — this isn't emerging technology, it's genuinely mature infrastructure.
It powers ML models inside key Microsoft products — Office, Azure, Bing — alongside dozens of community projects, a real signal of production battle-testing at scale.
The core idea, worth internalizing upfront: train anywhere, export to the portable ONNX format once, then let ORT's hardware-specific execution layer handle the rest — regardless of whether "the rest" means a cloud GPU cluster or a phone.
Execution Providers: The Actual Mechanism Behind "Any Hardware"
This is genuinely the concept that makes ONNX Runtime's pitch work, and it's worth understanding properly rather than treating as magic.
Execution Providers (EPs) are ONNX Runtime's extensible interface to different hardware acceleration libraries — CUDA, TensorRT, DirectML, OpenVINO, CoreML, QNN, and more. The same ONNX Runtime API sits on top of all of them, giving your application code a consistent interface regardless of which EP actually ends up executing a given operation.
ORT works with each EP through a GetCapability() interface, allocating specific nodes or sub-graphs of your model to whichever EP can handle them best on the current hardware.
EPs are registered in priority order — for example, ['CUDAExecutionProvider', 'CPUExecutionProvider'] means: try CUDA first for any node that supports it, and fall back to plain CPU execution for anything that doesn't.
This architecture abstracts away the hardware-specific library details entirely from your application code — you write against ORT's API once, and the actual execution backend is a runtime configuration choice, not something baked into your code.
This is genuinely the same abstraction principle from the local LLM tools comparison earlier in this series — recall llama.cpp's enormous backend list covering CUDA, ROCm, Metal, Vulkan, and more. ONNX Runtime is doing the identical job at the application-framework level rather than the raw inference-engine level.
Getting Started: Converting a Model to ONNX
The premise is genuinely simple: get a model from any framework that supports ONNX export, convert it, then run it through ORT.
import torch
torch.onnx.export(
model,
dummy_input,
"model.onnx",
input_names=["input"],
output_names=["output"],
dynamic_axes={"input": {0: "batch_size"}}
)
That dynamic_axes parameter matters more than it looks — without it, your exported model is locked to a fixed batch size, which genuinely limits deployment flexibility. Marking the batch dimension as dynamic lets the same exported file handle single-image inference and batched inference equally well.
Running Inference: The Same Pattern, Yet Again
If you've followed this series through the Raspberry Pi and LiteRT tutorials, this structure will feel genuinely familiar.
import onnxruntime as ort
import numpy as np
session = ort.InferenceSession("model.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
input_name = session.get_inputs()[0].name
output_name = session.get_outputs()[0].name
input_data = np.random.randn(1, 3, 224, 224).astype(np.float32)
result = session.run([output_name], {input_name: input_data})
Notice the providers list in that InferenceSession call — this is where you set the priority order discussed above. Pass a list with your preferred hardware first and CPU as the fallback, and ORT handles the actual routing decision per-operation automatically.
Choosing the Right Execution Provider for Your Deployment
Different EPs genuinely suit different deployment targets, and picking correctly matters for real performance.
CUDAExecutionProvider — standard GPU acceleration for NVIDIA hardware; delivers roughly 1.5-3x speedup over native PyTorch inference for typical computer vision and NLP models. If you're deploying on NVIDIA hardware, a GPU like the RTX 5070 or RTX 5080 gives you solid acceleration.
TensorRTExecutionProvider — squeezes further performance specifically from NVIDIA GPUs through additional graph optimization; note that TensorRT engine compilation happens on the first inference and can take 1-5 minutes — budget for that warm-up cost in latency-sensitive deployments.
DirectML — Windows-specific acceleration working across a broad range of GPU vendors, useful for Windows desktop apps that need hardware acceleration without committing to CUDA specifically.
CoreML — Apple's native acceleration path, the direct equivalent of what LiteRT uses for iOS deployment.
OpenVINO — Intel's acceleration path, particularly relevant for Intel CPU and integrated GPU deployments.
QNN — Qualcomm's NPU acceleration, the ORT-side equivalent of the vendor-specific NPU tooling discussed in the local LLM tools comparison.
This EP list is genuinely broader than LiteRT's hardware coverage — ONNX Runtime explicitly targets server and desktop GPU scenarios that LiteRT, with its mobile-and-embedded focus, doesn't prioritize as heavily.
ONNX Runtime Mobile: The Direct LiteRT Competitor
Worth naming specifically since this series just covered LiteRT in depth: ONNX Runtime Mobile is a real, actively maintained alternative for phone and edge deployment, not just a server-side tool awkwardly ported down.
ORT runs on every major platform including Windows, with ONNX Runtime Mobile specifically optimized for phones and edge devices.
The tradeoff versus LiteRT is genuinely about ecosystem breadth versus mobile-specific polish — LiteRT's automated hardware selection and GenAI-specific tooling are more mature for pure mobile use cases, while ORT's execution provider diversity gives it a genuine edge on Windows and server hardware.
Model support differs meaningfully too: ORT supports any model convertible to ONNX format — covering virtually every ML framework — while requiring that conversion step as genuine overhead LiteRT's more direct TensorFlow/Keras/PyTorch pipeline sometimes avoids.
Neither is strictly better — if your deployment spans server, desktop, and mobile with a genuinely diverse hardware mix, ORT's execution provider breadth is the stronger fit. If you're purely mobile-focused with GenAI ambitions, LiteRT's more specialized tooling has an edge.
Where ONNX Fits Across Frameworks You've Already Trained In
This matters directly for anyone who worked through the RL series earlier — your Stable-Baselines3 PyTorch policies, your Sentence Transformers embeddings, your scikit-learn preprocessing pipelines all export to the same universal format.
# A trained scikit-learn model
from skl2onnx import convert_sklearn
from skl2onnx.common.data_types import FloatTensorType
onnx_model = convert_sklearn(sklearn_model, initial_types=[("input", FloatTensorType([None, 4]))])
This is genuinely useful if your production stack mixes model types — a scikit-learn preprocessing step feeding into a PyTorch neural network, both exported to ONNX and run through the same InferenceSession interface, rather than juggling two entirely separate inference stacks in production.
On-Device Training: A Feature Worth Knowing About
Beyond pure inference, ONNX Runtime supports on-device training — taking an already-deployed inference model and further training it locally, directly on the device it's running on.
This enables more personalized, privacy-respecting experiences — a model that adapts to a specific user's data without that data ever needing to leave the device for retraining.
This connects directly back to the edge AI privacy discussion from earlier in this series — the same "data never leaves the device" principle that motivated edge AI generally, extended here to cover the training step too, not just inference.
ONNX Runtime vs LiteRT: When to Use Which
Both are excellent deployment runtimes, but they serve different primary use cases:
| Factor | ONNX Runtime | LiteRT (TFLite) |
|---|---|---|
| Framework support | PyTorch, TF, scikit-learn, XGBoost | TF, PyTorch, JAX |
| Hardware breadth | Servers, desktops, mobile, edge | Primarily mobile/embedded |
| Mobile tooling | Good (ORT Mobile) | Excellent (automated HW selection) |
| Server GPU | Excellent (CUDA, TensorRT) | Limited |
| GenAI support | Good | Excellent (LLMs, diffusion) |
| Windows desktop | Excellent (DirectML) | Limited |
Choose ONNX Runtime when deploying across diverse hardware (server + desktop + mobile) or using scikit-learn/XGBoost models.
Choose LiteRT when deploying primarily to mobile with GenAI ambitions.
Common Mistakes People Make
Forgetting dynamic_axes during export. Without it, your model is locked to whatever batch size you used during the export call — genuinely limiting for real deployment flexibility.
Not setting a CPU fallback in the providers list. If your preferred EP (CUDA, TensorRT) isn't available on a given deployment target, omitting the CPU fallback means outright failure instead of graceful degradation.
Ignoring TensorRT's first-inference compilation cost. That 1-5 minute warm-up on first run can genuinely blindside a latency-sensitive service if you haven't accounted for it in your deployment strategy.
Choosing ONNX Runtime when LiteRT's mobile-specific tooling would serve better. If you're purely mobile-focused with GenAI ambitions, LiteRT's automated hardware selection is genuinely more specialized for that exact use case.
Treating framework export as an afterthought rather than validating it. Always run inference on the exported ONNX model and compare outputs against the original framework's results before trusting it in production — subtle operator mismatches during conversion do happen.
Where ONNX Fits in This Series
ONNX Runtime is the deployment counterpart to everything covered in this series. Your Arduino TinyML models could export to ONNX for cross-platform deployment. Your Raspberry Pi ML models could run through ORT instead of TFLite. Your quantized models from the previous article can export to ONNX with quantization baked in.
The local LLM tools comparison covered llama.cpp's backend abstraction — ONNX Runtime is doing the identical job at the application framework level, abstracting CUDA, TensorRT, DirectML, CoreML, and OpenVINO behind one consistent API.
Next Steps: Master Cross-Platform ML Deployment
Ready to deploy models across any hardware stack? Educative offers interactive courses on ONNX Runtime, model optimization, and building production ML systems — practice in real sandboxed environments and learn by building actual deployment pipelines.
Wrapping This Up
ONNX Runtime solves the "trained here, need to deploy there" problem that quietly underlies most real ML deployment work: a portable model format plus an Execution Provider architecture that lets the identical file run optimally across CUDA, TensorRT, DirectML, CoreML, OpenVINO, and more, all behind one consistent API. Where LiteRT specializes deeply in mobile and embedded, ONNX Runtime spans a genuinely broader hardware range, particularly on server and desktop hardware.
Remember to always set a CPU fallback in your execution provider priority list, account for TensorRT's warm-up compilation cost in latency-sensitive deployments, and validate your exported model's outputs against the original framework before trusting it in production. FYI, if you've trained anything in this series using PyTorch or scikit-learn, it's already export-ready for this exact pipeline — the "write once, run anywhere" pitch genuinely holds up here more than it does for most cross-platform claims in software :)
Now go take one of the RL policies or classification models from earlier in this series, export it to ONNX, and run it through both the CPU and GPU execution providers side by side. Watching the identical model file route through completely different hardware backends without any code changes is genuinely the clearest way to see this abstraction actually working.