Local Voice Assistants: Build a Privacy-First Alexa Alternative

September 29, 202615 min readSam Austin
Contents

Local voice assistant running fully offline privacy-first smart speaker alternative
Local voice assistant running fully offline privacy-first smart speaker alternative

Figure 1: The form factor you're replacing — same kitchen counter, none of the data center

Alexa listens, transcribes, and sends your kitchen conversation to a data center. Every "what's the weather" gets logged somewhere you don't control. If that trade-off bothers you, the good news is the fully local alternative is genuinely buildable now, not just a hobbyist fantasy, and I'll walk you through exactly how to put it together.

I run something close to this setup at home. It's not quite as polished as commercial assistants, but the trade-off — zero cloud dependency and zero data leaving my network — is worth the extra setup for me.

The Four Roles Every Voice Assistant Needs

Before touching any software, understand the shape of the problem. A local voice assistant is four roles running on your own hardware: capture and trigger (wake word), transcribe (speech-to-text), understand (intent parsing or an LLM), and respond (text-to-speech). Each piece runs offline, and a lightweight protocol glues them together.

That protocol is called Wyoming, and it's the backbone of this entire ecosystem. Wyoming connects external voice services to Home Assistant using a small protocol, letting Assist use a variety of local speech-to-text, text-to-speech, and wake-word-detection systems interchangeably.

This is the privacy version of the same on-device bet our Edge AI for beginners guide describes: compute moves to where the data already is.

The Stack You'll Actually Build

Here's the five-component stack that's become the de facto standard for this project in 2026:

ComponentRoleTool
Wake wordListens continuously, triggers on your keywordopenWakeWord
Speech-to-textConverts your speech to textWhisper
Intent/brainsUnderstands the request, decides what to doHome Assistant Assist, optionally backed by a local LLM
Text-to-speechSpeaks the response backPiper
GlueConnects everythingWyoming Protocol

This isn't a niche setup either. The Whisper, Piper, and Wyoming Protocol integrations are each used by 8.9% of all active Home Assistant installations, which represents a genuinely real installed base for a fully local voice stack, not just a handful of tinkerers.

Setting Up Wyoming Services

Each piece of the stack runs as its own small server, typically in a Docker container, exposing a Wyoming TCP port that Home Assistant connects to. A typical multi-container layout looks like this:

  • wyoming-whisper — runs the Whisper STT service, default port 10300
  • wyoming-piper — runs Piper TTS, default port 10200
  • wyoming-openwakeword — runs wake-word detection, default port 10400
  • ollama (optional) — runs your local LLM, HTTP API on port 11434
  • homeassistant — the core instance, connects to all of the above

Wyoming services can run on another device on your local network, which matters if your Home Assistant box is a low-power Raspberry Pi but you've got a beefier machine elsewhere for the heavier Whisper or Ollama workloads. You're not locked into running everything on one box.

If hardware selection is still an open question, our mini PC and single-board computer roundup covers what actually fits an always-on inference box.

A Quick Note on Piper's Status

Worth flagging honestly: Piper TTS was archived on GitHub on October 6, 2025 and is now read-only. That sounds alarming, but it still works perfectly fine as a Home Assistant Wyoming add-on, and the community-maintained wyoming-piper wrapper around it continues to receive updates. Don't panic if you see "archived" on the repo, just know you're relying on a project that's functionally stable rather than actively evolving.

Setting Up Speech-to-Text With Whisper

Whisper handles the transcription piece, and you have two deployment options depending on your hardware:

pip install -U wyoming-faster-whisper

Or run it as a Docker container connected via the Wyoming protocol. The faster-whisper implementation is the practical choice for most home setups since it runs comfortably on CPU-only hardware without needing a dedicated GPU. Better resilience in isolated or offline environments is one of the real advantages of keeping this entirely local rather than routing through a cloud STT API.

If you'd rather go deeper on the C++ route than the Python wrapper, our Whisper.cpp on-device speech guide covers the lower-overhead alternative.

Setting Up Text-to-Speech With Piper

Piper handles the spoken response, and setup is straightforward:

docker run -v /path/to/local/data:/data \
    rhasspy/wyoming-piper \
    --voice en_US-lessac-medium

You'll pick a voice model — a .onnx file plus its config — and Piper serves it over the Wyoming protocol. There's an active community sharing custom voices too. FYI, if you want something more distinctive than the default voices, community-shared options exist on Hugging Face, including a fair number of novelty voices people have trained for fun.

A couple of practical notes: use --local-files-only once your voice model is cached to guarantee zero network calls, and if you're adding custom voices manually, drop them in the correct shared directory for your installation type, since Home Assistant's official add-on and a plain Docker container expect different paths.

Setting Up Wake Word Detection

openWakeWord handles the "hey assistant" trigger that starts the whole pipeline listening. It runs continuously and locally, watching for your chosen wake phrase without sending anything anywhere until triggered.

docker run rhasspy/wyoming-openwakeword

Community tooling has made this even easier to customize. A HACS-installable wakeword installer can automatically download and install .tflite wake-word files from GitHub repositories directly into Home Assistant's /share/openwakeword/ directory, so building or swapping a custom wake word doesn't require manual file wrangling.

Wiring It All Together in Home Assistant

Once your Wyoming services are running, connect them through Home Assistant's settings:

  1. Add the Wyoming Protocol integration for each service (Whisper, Piper, openWakeWord), either through auto-discovery or by entering the hostname and port manually
  2. Go to Settings → Voice Assistants and create a new Assistant pipeline
  3. Select your Whisper instance as the speech-to-text engine
  4. Select your Piper instance as the text-to-speech engine
  5. Select your wake word from openWakeWord
  6. Choose your conversation agent: Home Assistant's built-in Assist intents, or a local LLM if you want more natural understanding

Ever asked a smart speaker something slightly off-script and gotten a useless "I don't understand that"? That's usually the intent engine's limitation, not the speech recognition. This is exactly where the next piece comes in.

Adding a Local LLM as the Brain

Home Assistant's built-in Assist intent engine handles structured commands well ("turn on the kitchen lights"), but it's not built for open-ended conversation. For that, wire in a local LLM through Ollama as the conversation agent instead of, or alongside, Assist's intent matching.

docker run -d -p 11434:11434 ollama/ollama
ollama pull llama3.2

Point Home Assistant's conversation agent configuration at your Ollama instance's HTTP endpoint, and suddenly your assistant can handle genuinely conversational follow-ups rather than only rigid command phrases. A fully offline voice pipeline wires Whisper, Piper, and Ollama together through Home Assistant Assist this way, giving you speech-to-text, an LLM brain, and speech synthesis, with nothing leaving your network at any stage.

Which model you pull matters. Our local LLM tools comparison and Ollama setup guide both cover picking something small enough to answer in under a second — latency is the whole game in a voice pipeline.

IMO, start without the LLM. Get Whisper, Piper, and basic Assist intents working reliably first, then layer in Ollama once the core pipeline is solid. Debugging three moving pieces is hard enough before adding a fourth.

Hardware Considerations

You don't need exotic hardware for the base stack:

  • Whisper (base/small models): runs fine on CPU, a Raspberry Pi 5 8GB handles light usage
  • Piper: lightweight enough to run alongside Whisper on the same modest hardware
  • Ollama with a small local model: benefits significantly from a GPU or Apple Silicon, but smaller models (1B-3B range) work tolerably on CPU too
  • Wake word detection: extremely lightweight, runs continuously without meaningful resource strain

If you're running everything including an LLM on one box, plan for more headroom than the STT/TTS pieces alone would need. Splitting Ollama onto separate, more capable hardware and keeping Whisper/Piper on a lighter always-on device is a common and sensible split. Our Raspberry Pi ML guide covers what that board can realistically sustain.

A Practical Decision Framework

Let me save you some research time with straightforward guidance:

  1. First weekend? Wyoming add-ons for Whisper, Piper, and openWakeWord only — skip Ollama entirely
  2. Commands feel brittle? That's the intent engine, not the microphone; add Ollama as the conversation agent
  3. Responses too slow? Smaller Whisper model, smaller LLM, and separate the two onto different machines
  4. Pi running out of headroom? Move Whisper and Ollama to a mini PC, keep Home Assistant and the wake word on the Pi
  5. Want better raw transcription? The Whisper.cpp path gives you more control over the decode settings than the wrapper does

The mistake isn't picking the wrong row — it's starting with four moving parts and then wondering why none of them work.

Common Mistakes People Make

Forgetting --local-files-only on Piper

Recall the text-to-speech section directly — without it, the service may still reach out for model files it thinks it needs, which quietly breaks the privacy guarantee you built this for.

Voice not appearing after adding custom files

Recall the text-to-speech section directly — Home Assistant caches the voice list, so reload the integration rather than assuming something's broken.

Wrong directory for the add-on vs Docker

Recall the text-to-speech section directly — the official add-on expects /share/piper, a standalone Docker container expects its own mounted /data path, and only one of them is right for your install.

Jumping straight to an LLM brain

Recall the LLM section directly — conversational integration adds real complexity, and you want a known-good pipeline to compare it against.

  • Designing Voice User Interfaces by Cathy Pearl — the interaction design side nobody tells you about, covering turn-taking, error recovery, and what makes a voice flow feel natural rather than frustrating.
  • Speech and Language Processing by Daniel Jurafsky and James H. Martin — the reference behind the transcription and intent stages, for when you want to understand why the pipeline behaves the way it does.
  • AI at the Edge by Daniel Situnayake and Donna Dubinsky — the broader case for keeping inference local, and the practical constraints that shape every choice in this build.

Want to Go Deeper?

If you want structured practice on local and edge ML, Educative's ML courses include hands-on labs that pair well with this kind of self-hosted pipeline work. The unlimited plan is useful when you're working through several model types in one stretch.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

Is a local voice assistant really private?

Yes, if every stage runs on hardware you control. Wake word detection, speech-to-text, intent handling, and text-to-speech all execute on your own machines, so audio never leaves your network. The one discipline required is keeping voice models cached locally so no service quietly fetches a file it needs.

What is Wyoming Protocol?

Wyoming is a small protocol that connects external voice services to Home Assistant, letting it swap speech-to-text, text-to-speech, and wake-word systems interchangeably. Each service runs as its own lightweight server exposing a Wyoming TCP port, which is what makes the stack composable across machines.

What hardware do I need to run this?

Whisper's base and small models run comfortably on a Raspberry Pi 4 or 5 alongside Piper, and wake word detection is nearly free. The optional Ollama language model is the piece that benefits from a GPU or Apple Silicon, though 1B to 3B models are usable on CPU if you keep expectations modest.

Do I need a local LLM in the pipeline?

No. Home Assistant's built-in Assist intent engine handles structured commands such as turning lights on and off very reliably. Add Ollama only when you want open-ended conversational follow-ups, and add it after the base pipeline is stable rather than at the same time.

Is Piper text-to-speech still maintained?

Piper's repository was archived on GitHub in October 2025 and is now read-only, which sounds worse than it is. It still works as a Home Assistant Wyoming add-on, and the community wyoming-piper wrapper continues to receive updates, so you are relying on something stable rather than something abandoned.

Can the services run on different machines?

Yes. Wyoming services can run on any device on your local network, which is the sensible split when your Home Assistant instance is a low-power Raspberry Pi but you have a stronger machine elsewhere for Whisper or Ollama.

How does a local assistant compare to Alexa?

It is less polished out of the box and you assemble it yourself instead of unboxing a finished product. In exchange you get zero cloud dependency, no account requirement, and full control over what happens to your audio. Most people find that trade-off worth the setup time.

Wrapping This Up

A genuinely private, fully offline voice assistant isn't science fiction anymore. Wyoming Protocol glues together openWakeWord, Whisper, Piper, and Home Assistant Assist, with an optional local LLM through Ollama for more natural conversation, and none of it phones home.

Will it match Alexa's polish out of the box? Not quite — you're assembling this yourself instead of unboxing a finished product. But once it's running, and your kitchen conversations genuinely stay on your own network, that trade-off feels like a pretty easy one to make :)

When the LLM stage starts feeling sluggish, our LM Studio tutorial covers a friendlier local serving path, and the local LLM tools comparison weighs the alternatives.