Contents
Every voice assistant demo you've ever been impressed by owes something to a single insight: OpenAI trained Whisper on 680,000 hours of messy, imperfect, internet-scraped audio-transcript pairs instead of a small curated dataset — and that "weakly supervised" approach at scale produced a genuinely more robust speech recognizer than the carefully labeled datasets that came before it. whisper.cpp is what happens when someone takes that model and strips it down to run entirely offline, on hardware as small as a Raspberry Pi.
This is genuinely the audio counterpart to everything the local LLM arc of this series covered — same philosophy (own the weights, own the compute, own the privacy), same core contributor even — Georgi Gerganov, the person behind llama.cpp and the GGUF format, wrote this too. If GGUF's structure made sense to you from that earlier article, whisper.cpp's model files will feel genuinely familiar.
By the end of this guide, you'll have whisper.cpp built and transcribing real audio entirely offline, understand which model size actually fits your hardware, and know when to reach for one of its more specialized alternatives instead. IMO, watching accurate transcription happen with your wifi turned off is a genuinely similar thrill to the first local LLM moment from earlier in this series :)
What Whisper Actually Is, Briefly
Whisper is OpenAI's general-purpose automatic speech recognition model, released September 2022 under the MIT license — genuinely open, weights included, no per-minute fee for running it yourself. It uses an encoder-decoder Transformer architecture, converting raw audio into log-mel spectrograms, then autoregressively decoding text — supporting transcription across 99 languages, translation to English, language identification, and timestamp prediction, all within one unified model.
The MIT license on both code and weights is exactly why an ecosystem of ports like whisper.cpp exists at all — the same open-weights philosophy that let llama.cpp exist for LLMs applies here.
What whisper.cpp Adds on Top
whisper.cpp is a lightweight C/C++ reimplementation of Whisper, focused specifically on efficient on-device inference — genuinely the same architectural philosophy as llama.cpp, applied to a different model family entirely.
- Pure C/C++ with minimal runtime dependencies — no PyTorch, no Python interpreter required at inference time, just a compiled binary.
- Runs across an enormous hardware range — from Raspberry Pi to Apple Silicon to desktop GPUs — with multiple acceleration backends including CUDA, Vulkan, CoreML, OpenVINO, and Moore Threads.
- Uses GGML weights with quantized model support — recall the GGUF article from earlier in this series; whisper.cpp uses the predecessor GGML format specifically, packaging Whisper's weights the same conceptual way, just for this model family.
- On Apple Silicon specifically, tensor operations are heavily optimized — using ARM NEON SIMD instructions or the Accelerate framework's AMX coprocessor for larger computations, genuinely squeezing real performance out of the exact hardware most people already own.
Building whisper.cpp
git clone https://github.com/ggml-org/whisper.cpp.git
cd whisper.cpp
cmake -B build
cmake --build build --config Release
Notice this is genuinely the identical CMake build pattern from the llama.cpp tutorial earlier in this series — same maintainer ecosystem, same build philosophy. If you built llama.cpp from source already, nothing here should feel unfamiliar.
Downloading a Model
bash ./models/download-ggml-model.sh base.en
Whisper ships in a genuine range of sizes, and picking correctly matters enormously for whether this runs comfortably on your target hardware.
| Model | Size | Best For |
|---|---|---|
| tiny / tiny.en | ~75MB | Microcontroller-adjacent hardware, extreme speed |
| base / base.en | ~142MB | Raspberry Pi and similar constrained hardware |
| small / small.en | ~466MB | Balanced accuracy and speed on modest desktop hardware |
| medium | ~1.5GB | Higher accuracy, genuinely capable hardware needed |
| large-v3 | ~2.9GB | Maximum accuracy, desktop/server-class hardware |
The .en suffix models are English-only variants, genuinely worth preferring if your use case doesn't need multilingual support — they're smaller and faster for equivalent accuracy on English audio specifically.
Your First Transcription
./build/bin/whisper-cli -m models/ggml-base.en.bin -f samples/jfk.wav
This produces timestamped transcription output directly to your terminal — genuinely the whole pipeline in one command, no server, no API call, nothing leaving your machine.
Running Quantized Models for Constrained Hardware
Just like GGUF's Q4_K_M tags from earlier in this series, whisper.cpp models come in quantized variants for tighter memory budgets.
bash ./models/download-ggml-model.sh tiny.en-q5_1
This is exactly the same quantization principle from the model compression articles — trading a small amount of accuracy for a meaningfully smaller model, genuinely useful when targeting Raspberry Pi-tier hardware rather than a desktop.
Real-Time Streaming Transcription
Static file transcription is a good starting point, but the genuinely more useful project is live microphone transcription.
./build/bin/whisper-stream -m models/ggml-base.en.bin -t 4 --step 500 --length 5000
That --step and --length pairing controls the streaming window — how often the model processes a new audio chunk versus how much context it considers at once. Getting this balance right genuinely matters: too short a step produces choppy, disconnected transcription; too long introduces noticeable lag before words appear.
If you're running streaming transcription on consumer hardware, a capable GPU helps significantly — the RTX 5070 and RTX 5080 both handle real-time Whisper inference with headroom to spare, and the faster feedback loops genuinely matter when tuning streaming parameters.
Where whisper.cpp Fits Against Its Alternatives
Worth being honest here, since newer options genuinely compete on specific axes — the same "pick the right tool for your actual constraint" principle that's run through this whole series.
- faster-whisper (CTranslate2-based) is genuinely the fastest path specifically on NVIDIA hardware — if you're deploying on a CUDA-capable GPU rather than Apple Silicon or a Raspberry Pi, this often edges out whisper.cpp on raw throughput.
- WhisperX wraps Whisper with phoneme-level alignment and speaker diarization — reach for this specifically when you need to know who said something and exactly when, not just what was said.
- distil-whisper is a genuinely distilled model (recall the knowledge distillation article) — roughly 6x faster and about 50% smaller than the original, within 1% word error rate on long-form audio. This is exactly the distillation technique from that earlier article applied concretely to speech recognition.
- Argmax's WhisperKit is built by former Apple engineers with deeper Neural Engine optimization specifically for Apple platforms — worth considering if you're building an Apple-exclusive product and want to squeeze out the last bit of hardware-specific performance whisper.cpp's more general approach doesn't fully capture.
- NVIDIA's Parakeet is a genuinely competitive newer alternative worth knowing about, licensed CC-BY-4.0 rather than MIT, and increasingly benchmarked directly against Whisper for accuracy.
The practical guidance, mirroring the local LLM tools comparison from earlier in this series: pick your runtime by your actual hardware, not by hype. whisper.cpp and MLX for Apple Silicon, faster-whisper for NVIDIA, WhisperX when you need speaker labels, distil-whisper when raw speed matters more than the last percentage point of accuracy.
Realistic Performance Expectations
- Real-time transcription is genuinely achievable on modest hardware with the smaller model variants — this isn't an aspirational claim, it's the actual documented experience across Raspberry Pi through desktop deployments.
- On Apple Silicon specifically, Metal and CoreML acceleration deliver real-time transcription speeds even with larger models, thanks to the hardware-specific optimization whisper.cpp includes for that platform.
- Word error rates under 6% are achievable with modern runtimes and appropriately sized models on clear audio — genuinely usable accuracy for real applications, not just demo-quality results.
- This exact pipeline already powers real dictation tools across macOS, Windows, and Linux — raw audio never leaves the device, and model size is the dial you turn to trade speed against accuracy for your specific hardware.
Combining With the Local LLM Stack
This is genuinely worth connecting explicitly to the Ollama and local RAG chatbot articles from earlier in this series: a voice-driven local assistant is just whisper.cpp feeding transcribed text into a local LLM, with zero cloud dependency anywhere in the pipeline.
Microphone → whisper.cpp (speech-to-text) → Ollama (local LLM) → response
Every piece of that pipeline has already been covered in this series individually — whisper.cpp handles the speech-to-text step, and everything from the Ollama setup guide and local RAG chatbot articles handles what happens with the transcribed text afterward. This is genuinely just the assembly step, the same way the local RAG article was an assembly of pieces you already understood.
Common Mistakes People Make
- Downloading the largest model available by default. Match model size to your actual hardware and accuracy needs —
tiny.enorbase.engenuinely suffice for many real-time applications on constrained devices. - Using a multilingual model when your use case is English-only. The
.envariants are smaller and faster for equivalent accuracy on English audio specifically — don't pay the multilingual overhead if you don't need it. - Choosing whisper.cpp when your actual hardware is an NVIDIA GPU. faster-whisper genuinely outperforms it there — whisper.cpp's strength is its broad cross-platform reach, particularly Apple Silicon, not necessarily peak NVIDIA throughput.
- Ignoring the streaming step/length tuning. Default values won't necessarily suit your specific latency-versus-coherence tradeoff — this genuinely needs adjustment for a smooth live-transcription experience.
- Expecting speaker diarization from whisper.cpp alone. It focuses exclusively on transcription — reach for WhisperX specifically when you need to know who said what.
Wrapping This Up
whisper.cpp brings the exact same local-first philosophy from this series' LLM tooling arc to speech recognition: a lightweight C/C++ port of an open-weights model, quantized and optimized for hardware ranging from Raspberry Pi to Apple Silicon, transcribing audio with zero data ever leaving the device. Choosing the right model size and runtime for your specific hardware matters here exactly as much as it did for local LLMs.
Remember to match model size to your actual hardware and accuracy needs, and know that faster-whisper, WhisperX, and distil-whisper each exist to solve a specific, different constraint whisper.cpp's general-purpose design doesn't optimize for by default. FYI, the fact that the same person behind llama.cpp and GGUF also built this genuinely isn't a coincidence — it's the same "strip inference down to portable C/C++, quantize the weights, run anywhere" philosophy applied a second time to a completely different model family :)
Now go pipe whisper.cpp's transcription output directly into the Ollama-based local RAG chatbot from earlier in this series and ask it a question out loud instead of typing it. That's genuinely the moment this whole local-AI series stops being separate tutorials and starts feeling like one coherent, fully offline stack you built piece by piece.