Contents
Figure 1: The app that keeps working when the Wi-Fi doesn't
Your user is on a plane, in a basement server room, or just has spotty Wi-Fi, and your "AI-powered" app goes completely dark the moment the network drops. That's a solvable problem now, and not just for native mobile apps. Even browser-based web apps can run real ML models fully offline today. Let's build one properly.
I've shipped a small offline classification feature into a web app before, and the architecture surprised me with how little server infrastructure it actually needed once the model was cached client-side.
The Core Idea: Models Live on the Device
The fundamental shift here is simple to state and genuinely powerful in practice: the inference runtime lives inside your app, not on a server your user has to reach. Once the model weights download and cache once, your app works completely offline afterward — no round trip, no dependency on connectivity at all.
Ever built a feature that felt broken every time someone's Wi-Fi hiccupped? That entire category of bug disappears once inference runs locally.
This is the browser-side version of the same idea behind our Edge AI beginner's guide: push computation toward the user instead of pulling them toward your server.
The Browser Architecture, Piece by Piece
For web apps specifically, the architecture has a consistent shape across every library you'll use:
User's browser
├── Your app code (React / Vanilla / whatever)
├── ML library (Transformers.js / WebLLM / ONNX Runtime)
│ ├── WebGPU backend ──► User's GPU (fast, requires support)
│ └── WASM backend ──► User's CPU (slower, works everywhere)
└── Cache API / IndexedDB
└── Model files (cached after first download, never re-sent)
Two runtime paths matter here. WebGPU is the fast path, running inference directly on the user's GPU when their browser supports it. WASM is the fallback, a compiled ML runtime (usually a port of ONNX Runtime or llama.cpp) running on CPU, slower but universally supported. Most production setups use both, attempting WebGPU first and falling back gracefully.
Model files get cached in the browser after the first download, so repeated use, including fully offline use, doesn't require re-fetching anything. That first-load download is the only moment your app genuinely needs a network connection.
Picking Your Library
Three main libraries dominate this space, and they're not interchangeable — each targets a different job:
| Library | Best For | Underlying Tech |
|---|---|---|
| WebLLM | Chatbot-style LLM inference | Built on Apache TVM, currently the gold standard for browser chatbots |
| Transformers.js | Broad tasks: classification, vision, speech, embeddings | Wraps ONNX Runtime, falls back to WASM automatically |
| ONNX Runtime Web | Custom or non-Hugging-Face models exported to ONNX | Direct ONNX standard support, framework-agnostic |
WebLLM is optimized specifically for LLMs, and as of late 2026 it offers a familiar, OpenAI-compatible API, so switching code from a cloud LLM call to local inference is a minimal-diff change:
import * as webllm from "@mlc-ai/web-llm";
const engine = await webllm.CreateMLCEngine(
"Llama-3-8B-Instruct-v0.1-q4f16_1-MLC",
{ initProgressCallback: (report) => console.log(report.progress) }
);
Transformers.js is the broader tool. Developed by Hugging Face, it covers vision, embeddings, and speech recognition, not just text generation, acting as an ONNX Runtime wrapper with intelligent WASM fallback:
import { pipeline } from '@xenova/transformers';
const classifier = await pipeline('sentiment-analysis');
const result = await classifier("This offline app actually works!");
ONNX Runtime Web is the right pick when your model doesn't come from the Hugging Face ecosystem at all. It's a flexible engine for running ONNX-standard models directly, and since ONNX is framework-agnostic, you can export from PyTorch, TensorFlow, Keras, or scikit-learn and run the result client-side. ONNX models are widely represented on Hugging Face too, so this is a solid default when you want more control over a customized model — our ONNX Runtime tutorial covers that export step in depth.
IMO, most projects should start with Transformers.js unless you specifically need chatbot-grade LLM performance, since it covers more task types and has the gentlest learning curve.
A Real Combined Example
Production apps commonly combine libraries for different jobs within the same app. One pattern: WebLLM handles reasoning and conversation, while Transformers.js handles specialized tasks like speech-to-text or sentiment scoring, each library doing what it's actually best at rather than forcing one tool to cover everything:
// Reasoning: WebLLM
const chatEngine = await webllm.CreateMLCEngine("Llama-3-8B-Instruct-v0.1-q4f16_1-MLC");
// Specialized task: Transformers.js
const whisper = await pipeline('automatic-speech-recognition', 'Xenova/whisper-tiny.en');
const sentiment = await pipeline('sentiment-analysis');
This split matters because a general chat LLM is overkill and often slower for a narrow task like sentiment scoring, while a small specialized model handles it near-instantly. The Whisper line above is the same shape as our on-device speech recognition walkthrough, just moved from a native runtime into a browser tab.
Making Your App Actually Work Offline
Downloading models client-side gets you partway there. To make the app genuinely usable with no connection at all, you need standard offline-web patterns layered on top:
- Service worker + PWA manifest: Caches your app shell (HTML, CSS, JS) so the app loads at all without network
- Cache API for model weights: Browsers cache large model files automatically after first fetch, but explicit cache management gives you more control over eviction and versioning
- Graceful first-load messaging: Tell users clearly that the first launch needs a connection to download models, then everything after works offline
- IndexedDB for user data: Store conversation history or app state locally so it survives without a backend
None of this is exotic web development. It's the same offline-first PWA toolkit that's existed for years, just paired with an inference runtime instead of static content.
Checking Support Before You Commit
Not every browser supports WebGPU yet. Current baseline support includes Chrome, Edge, and recent Firefox and Safari versions, but you should always feature-detect rather than assume:
if (navigator.gpu) {
// Attempt WebGPU path
} else {
// Fall back to WASM
}
Libraries like Transformers.js handle this fallback logic internally in most cases, but knowing which path your users are actually landing on matters when you're debugging a performance complaint. A user reporting "it's really slow" might simply be running on the WASM fallback because their browser or hardware doesn't support WebGPU.
What This Doesn't Cover: Native Apps
Everything above is for web apps specifically. If you're building a native mobile or desktop app, the same underlying idea — ship the runtime with the app instead of relying on a server — applies but through different tools: CoreML/ANE and MLX on Apple platforms, LiteRT on Android, ExecuTorch across both, and native llama.cpp for desktop.
The architecture principle is identical even though the specific runtime differs by platform. Our Core ML tutorial, Apple Neural Engine explainer, and TensorFlow Lite guide each cover one of those targets. Some cross-platform SDKs now abstract this entirely — picking the right backend, WebLLM in a browser context, native llama.cpp in an Electron or mobile shell, automatically based on what the device supports — so your application code stays the same regardless of where it's running.
A Practical Decision Framework
Let me save you some research time with straightforward guidance:
- Mostly text classification, embeddings, or vision? Start with Transformers.js and a small ONNX model — one dependency, gentlest learning curve
- Building a chat interface in the browser? WebLLM for the OpenAI-compatible API and real chatbot performance
- Have your own trained model already? Export to ONNX and run it through ONNX Runtime Web for full control
- Speech-to-text offline? Whisper-tiny via Transformers.js, near-instant on CPU and a comfortable WebGPU win
- Shipping to mobile too? Look at MediaPipe or Edge Impulse for the native side of the same architecture
- Need a chat model offline on desktop? Our local LLM tools comparison covers Ollama and friends, which pair with this approach
The mistake isn't picking the wrong library — it's shipping an 8B model to first-time visitors without telling them what's coming.
Common Mistakes People Make
Assuming WebGPU everywhere
Recall the feature-detect section — test on a browser or device that forces the CPU path before you blame your model for being slow.
Forgetting the first-load story
Recall the offline section — an unexplained multi-hundred-MB download on first visit will confuse and frustrate users if you don't set expectations.
Not versioning cached models
Recall the Cache API notes — when you update a model, make sure old cached weights don't silently linger and get used instead.
Oversizing the first model
Recall the library discussion — a small model under 1GB makes first load painless; an 8B-parameter LLM means a genuinely long initial download users need to be warned about.
Skipping the offline test
Recall the support section — kill your Wi-Fi before shipping, not after someone files an issue about it.
Recommended Books
- Progressive Web Apps by Tal Ater — the service worker and Cache API half of this article, explained properly, which is exactly what turns a model download into a genuinely offline app.
- High Performance Browser Networking by Ilya Grigorik — how the browser actually moves bytes, why first-load cost feels so much larger than it should, and where caching decisions pay off.
- AI and Machine Learning for On-Device Development by Laurence Moroney — the native counterpart to everything above, for when your offline requirement moves out of the browser.
Want to Go Deeper?
If you want structured practice on ML systems and deployment, Educative's ML courses include hands-on labs that pair well with this kind of work. The unlimited plan is useful when you're working through several targets in one stretch.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
Can AI apps really work fully offline?
Yes. Once the model weights download and cache on the device, inference runs entirely locally with no server round trip. The only moment your app genuinely needs a network connection is the first load, when the model file is fetched.
What is WebGPU and why does it matter for offline AI?
WebGPU is the fast path: it runs inference directly on the user's GPU from the browser. It is currently supported in Chrome, Edge, and recent Firefox and Safari versions, but you should feature-detect with navigator.gpu and fall back to the WASM path on CPU rather than assume it exists.
Which library should I use for offline ML in the browser?
Most projects should start with Transformers.js, since it covers classification, vision, speech, and embeddings with automatic WASM fallback. Use WebLLM specifically for chatbot-grade LLM inference, and ONNX Runtime Web when you have a custom model exported to ONNX that is not part of the Hugging Face ecosystem.
How big are offline models?
It varies widely. A small classification or Whisper-tiny model can sit well under 1GB and makes first load painless, while an 8B-parameter LLM means a genuinely long initial download. Pick your model size deliberately relative to your users' patience for that first visit.
Does the app need internet on first launch?
Yes, for the model download. Be explicit about it: an unexplained multi-hundred-MB transfer on first visit confuses and frustrates users. After that first fetch, cached weights are reused and the app works with no connection at all.
Do I need a native app to ship offline ML?
No. Everything in this article works in a browser tab with a service worker and cached model files. Native builds use different runtimes such as CoreML, LiteRT, ExecuTorch, or native llama.cpp, but the architecture principle is identical: ship the runtime with the app instead of relying on a server.
How do I test offline behaviour?
Force both paths. Disable your network and confirm the cached app still loads and infers, and separately force the WASM CPU path so you know what users without WebGPU will experience. A performance complaint often turns out to be a user sitting on the fallback.
Wrapping This Up
Offline-capable AI apps are genuinely achievable in the browser today, no native app store deployment required. WebLLM handles conversational AI, Transformers.js covers the broader task landscape, and ONNX Runtime Web gives you flexibility for custom models, all running through WebGPU when available and falling back to WASM everywhere else.
Will every model fit comfortably in a browser tab? No, and picking model size relative to your users' patience for a first download matters more than almost anything else here. But once that download happens once, your app keeps working on a plane, in a basement, anywhere. Build a small offline classifier tonight with Transformers.js, kill your Wi-Fi, and watch it keep working anyway.
When it's time to extend that same thinking past the browser, our local voice assistant guide and Edge AI beginner's guide pick up right where this leaves off.