Sam Austin AI

llama.cpp Tutorial: Run Large Language Models on Your CPU (2026)

September 7, 2026 14 min read Sam Austin
Contents
llama.cpp Tutorial Run Large Language Models on CPU GGUF Quantization Guide
llama.cpp Tutorial Run Large Language Models on CPU GGUF Quantization Guide

Figure 1: Building and running llama.cpp — the engine powering most local LLM tools

Ollama makes running local LLMs feel like magic — one command, done. llama.cpp is what happens when you actually want to see the trick. No package manager abstraction, no background service hiding the details — just the raw C/C++ engine that Ollama, LM Studio, and half the tools in this space are quietly built on top of.

I'm writing this as the "go one level deeper" follow-up to the Ollama setup guide — that article got you chatting with a local model in minutes. This one gets you understanding exactly what's happening when you do that, plus genuine control over quantization, backend selection, and performance tuning that Ollama's convenience layer intentionally hides from you.

By the end of this tutorial, you'll have built llama.cpp from source, run a quantized model directly, exposed your own local API server, and understand quantization well enough to pick the right tradeoff for your hardware yourself. IMO, there's real value in doing this the "hard way" once, even if you go back to Ollama for daily use afterward :)

What llama.cpp Actually Is

llama.cpp is a pure C/C++ implementation of LLM inference, created by Georgi Gerganov in March 2023 and now maintained by the ggml-org community. Its GitHub description is deliberately minimal — "LLM inference in C/C++" — and that simplicity is genuinely the entire point.

No heavyweight runtime dependencies like PyTorch or CUDA developer toolkits baked into the inference path — it compiles to a handful of small native binaries that run almost anywhere: Linux, macOS, Windows, Raspberry Pi, Android, even inside a browser via WebGPU.

It reads models in the GGUF format — the successor to the older GGML format — packing weights, tokenizer, and metadata into a single portable file.

GGUF supports aggressive quantization, shrinking 16-bit weights down to 8, 6, 5, 4, 3, or even 2 bits, letting a model that needs a data-center GPU at full precision instead fit on a gaming card or run entirely in CPU RAM.

That quantization story is exactly why this tutorial exists — understanding it properly is the actual payoff for building from source instead of just installing Ollama.

Building From Source

Pre-built binaries are available, but compiling yourself is the standard practice if you want to understand your specific hardware's requirements and squeeze out maximum tokens-per-second.

git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build --config Release

The build system has fully migrated to CMake, making dependency management genuinely straightforward on modern Ubuntu or Fedora systems. Make sure cmake is actually installed first — on a fresh Linux box or a cloud instance, you may need:

apt-get update
apt-get install -y cmake libcurl4-openssl-dev

Building With GPU Acceleration

If you want CUDA or Metal support rather than pure CPU inference, confirm the relevant toolkit is installed before configuring — the NVIDIA CUDA toolkit (12.x+) or Apple Xcode, depending on your platform. The build process detects and links against these automatically once they're present.

If you're planning to use GPU acceleration, a modern GPU like the RTX 5070 with 12GB VRAM handles most local LLM workloads comfortably. For larger models or faster throughput, the RTX 5080 with 24GB VRAM gives you serious headroom — and building with the right backend from the start avoids a rebuild later.

This is genuinely worth doing even on a CPU-first tutorial — llama.cpp was originally built CPU-first, but running with -ngl (GPU layer offload) whenever a GPU is available dramatically changes throughput.

The Core Tools You'll Actually Use

Four binaries matter most for everyday use, and it's worth knowing what each one is for before touching any of them.

llama-cli — interactive terminal chat, the most direct way to talk to a model.

llama-server — spins up an OpenAI-compatible HTTP API, the same interface pattern you saw with Ollama.

llama-quantize — compresses a full-precision GGUF file down to a smaller quantized version.

llama-bench — benchmarks throughput, useful for comparing quantization levels or backend configurations on your specific hardware.

Getting Your First Model Running

The fastest path skips manual conversion entirely — Hugging Face hosts pre-quantized GGUF files from trusted publishers, and llama.cpp can pull them directly.

./build/bin/llama-cli -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q4_K_M

Always download models from verified publishers like Bartowski or TheBloke on Hugging Face rather than random uploads — and if you download a file manually rather than through the -hf flag, verify its SHA256 checksum, since a corrupted GGUF file causes silent failures or segfaults during context loading rather than a clean error.

Starting a Local API Server

./build/bin/llama-server -hf bartowski/Llama-3.2-3B-Instruct-GGUF:Q4_K_M --host 0.0.0.0 --port 8080 -ngl 99

This exposes a standard chat completions endpoint compatible with most existing AI tooling — the same OpenAI-compatible pattern Ollama uses, just running directly off llama.cpp with no wrapper in between. That -ngl 99 flag offloads as many layers as possible to GPU; drop it (or set it to 0) for pure CPU inference.

Understanding Quantization Properly

This is genuinely the concept worth taking time with, since it's the entire reason llama.cpp exists as a project.

Quantization reduces the numeric precision of model weights — typically from 32-bit or 16-bit floats down to much smaller integer representations — which shrinks file size and speeds up inference, at some cost to accuracy, usually measured in perplexity or KL divergence.

Q4_K_M offers the best balance for most general engineering and chat tasks — a genuinely reasonable default when you're not sure what else to pick.

Q6_K is preferable for reasoning-heavy workloads specifically, where accuracy matters more than raw speed or file size.

Q8_0 sits close to full precision, useful when quality loss from quantization is genuinely unacceptable for your use case, at the cost of a much larger file.

The lower the quantization bit-width, the faster and smaller the model — but with correspondingly reduced accuracy. This tradeoff is the entire design space llama.cpp's quantization tooling exists to navigate.

Quantizing a Model Yourself

If you have a full-precision GGUF file and want to compress it yourself rather than downloading a pre-quantized version:

./build/bin/llama-quantize model-f32.gguf model-Q4_K_M.gguf Q4_K_M

A couple of flags are worth understanding before running this: --allow-requantize lets you requantize a tensor that's already been quantized, but this can severely reduce quality compared to quantizing directly from 16-bit or 32-bit — avoid it unless you have a specific reason. --leave-output-tensor skips requantizing the output layer specifically, sometimes preserving more quality at a small size cost.

You can minimize accuracy loss further using an imatrix file during quantization — if you genuinely need the best possible quality at a given bit-width, this is the lever to reach for rather than just bumping up to a higher bit-width across the board.

Skipping the Manual Process Entirely

If building your own quantized GGUF sounds like more setup than you want, Hugging Face's GGUF-my-repo space lets you build custom quants without any local setup at all, syncing from llama.cpp's main branch every six hours — genuinely the easiest path if you just need one specific quantization level someone hasn't already published.

Converting a Model to GGUF From Scratch

If you're starting from a model's original safetensors format rather than an already-converted GGUF, the pipeline has one extra step before quantization.

python convert_hf_to_gguf.py /path/to/model --outfile model-f32.gguf
./build/bin/llama-quantize model-f32.gguf model-Q4_K_M.gguf Q4_K_M

This full pipeline — download, convert, quantize — can genuinely take hours depending on model size, bandwidth, and system resources, well beyond typical tutorial reading time. If a pre-quantized GGUF already exists for your model of choice, skip this entirely and download that instead.

Benchmarking Your Setup

Once you're running inference, llama-bench tells you concretely how your specific hardware and quantization choice actually perform together, rather than trusting a generic benchmark from someone else's machine.

./build/bin/llama-bench -m model-Q4_K_M.gguf

This is genuinely worth running whenever you change quantization level or backend configuration — the honest way to know whether a change actually helped, rather than assuming based on the spec sheet alone.

Installing Without Building (If You Change Your Mind)

If you've gone through this whole build process and decide you'd rather use a package manager after all, that path exists too.

brew install llama.cpp        # macOS
winget install llama.cpp      # Windows

Both give you llama-server and llama-cli directly, without a manual build step — genuinely reasonable if you've now learned what's happening under the hood and just want the convenience going forward.

Common Mistakes People Make

Downloading GGUF files from unverified sources. A corrupted or maliciously modified file can cause silent failures or crashes — stick to established publishers like Bartowski or TheBloke, and verify checksums for anything downloaded manually.

Requantizing an already-quantized model without understanding the quality cost. --allow-requantize exists for convenience, but quantizing directly from a full-precision source produces meaningfully better results.

Assuming pre-built binaries are optimized for your hardware. They often lack backend-specific optimizations or miss support your specific GPU could use — compiling from source remains the standard practice for anyone chasing maximum throughput.

Picking a quantization level without benchmarking. Q4_K_M is a reasonable default, but llama-bench on your actual hardware is the only way to know if a different level genuinely serves your workload better.

Forgetting -ngl when a GPU is actually available. Without it, llama.cpp may default to leaving GPU compute unused even when your build supports it.

Where This Fits With the Rest of the Local LLM Stack

Recall from the local LLM tools comparison earlier in this series: Ollama and LM Studio both run llama.cpp underneath their own interfaces. Everything you just learned about quantization, GGUF, and the core binaries is directly transferable — when Ollama silently picks a quantization level for you, you now understand exactly what tradeoff it's making on your behalf.

Use llama.cpp directly when you want that control back — custom quantization, specific backend tuning, or deployment onto genuinely unusual hardware (Raspberry Pi, Android, embedded environments) where Ollama's abstraction doesn't fit as cleanly.

Next Steps: Deepen Your LLM Knowledge

Ready to go beyond running models? Educative offers interactive courses on LLM fundamentals, transformer architectures, and building production AI applications — learn by coding in sandboxed environments rather than just reading tutorials.

Wrapping This Up

llama.cpp strips local LLM inference down to its actual mechanics: a portable GGUF file, a quantization level chosen deliberately rather than automatically, and a handful of small native binaries doing the real work. Building from source and running llama-quantize and llama-bench yourself turns "local AI just works" into "I understand exactly why, and I can tune it."

Remember that Q4_K_M is a solid general default while Q6_K suits reasoning-heavy workloads better, and that verifying your GGUF source and checksum matters — a corrupted file fails silently rather than with a clear error. FYI, everything you learned here transfers directly the next time you use Ollama or LM Studio, since both are running this exact engine underneath their friendlier interfaces :)

Now go quantize the same model at two different bit-widths and run them both through llama-bench side by side. That direct comparison is genuinely the fastest way to build real intuition for the speed-versus-quality tradeoff this whole tutorial has been building toward.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles