Benchmarking Local LLMs: Tokens per Second Across Hardware

September 29, 202613 min readSam Austin
Contents

Benchmarking local LLM tokens per second across GPU Apple Silicon and CPU hardware
Benchmarking local LLM tokens per second across GPU Apple Silicon and CPU hardware

Figure 1: The hardware question, answered with a stopwatch instead of a spec sheet

"How fast will this model run on my machine?" is the question everyone asks before buying hardware, and it's also the question most badly answered online. Marketing specs tell you peak theoretical throughput. Real benchmarks tell you what you'll actually feel while chatting. Let's get you real numbers and a repeatable way to measure your own setup.

I've burned money on hardware based on vague online claims before. This time, let's use actual measured figures and tell you exactly how to reproduce them yourself.

Why Tokens Per Second Is the Metric That Matters

Real tokens-per-second benchmarks are the only metric that tells you whether your local LLM will feel instant or painful. Not parameter count, not VRAM capacity alone, not a marketing TOPS number. Actual measured generation speed, under conditions matching how you'll really use it.

Two numbers matter together:

  • Time to first token (TTFT): how long before the response starts appearing
  • Tokens per second (TPS): how fast text streams out once generation starts

A model with slow TTFT but fast TPS feels sluggish to start then snappy. A model with fast TTFT but slow TPS feels responsive at first, then drags. Benchmark both, not just one.

Real Numbers Across Hardware Tiers

Here's what's actually been measured, using llama.cpp with Q4_K_M quantization (Q2_K for 70B models on 24GB cards), single batch, 2048 context, warm model:

Hardware7B Model70B Model (quantized)
RTX 4090~135 tok/s~18 tok/s (Q2)
RTX 3090~95 tok/s~10 tok/s (Q2)
RTX 3060 12GB~45 tok/sNot practical
Apple M3 Max 64GBSolid~5 tok/s
CPU-onlyUnder 10 tok/sPainful

That gap between the RTX 4090 and CPU-only is roughly 13x on a small model. On a 70B model, even the 4090 slows down considerably once quantized down to fit its 24GB, which tells you something important: model size relative to VRAM matters more than raw GPU horsepower once you're memory-bound.

If you're shopping from these numbers, our GPU guide for running local LLMs breaks the card choices down by model size.

Apple Silicon Numbers

Apple's unified memory architecture changes the calculus. An M4 Pro with 48GB of RAM can run Qwen 3 32B at Q4 quantization at roughly 15 to 22 tokens per second, and an M4 Max with 128GB of unified memory can handle 70B models that simply wouldn't fit on a single consumer GPU at all. That's the trade-off: Apple Silicon loses on raw bandwidth — an RTX 5090's 1,792 GB/s bandwidth far outpaces most M-series chips — but wins on being able to fit bigger models in the first place.

The High-End Reality Check

At the top end, an RTX 5090 running llama.cpp or LM Studio hit roughly 240 tokens per second decoding on an 18.63GB model that fits entirely in its 32GB of GDDR7 memory. Compare that against unified-memory platforms like NVIDIA's GB10, which measured around 44 tokens per second on the same workload — more than 5x slower on raw single-stream decode, precisely because of that bandwidth gap (1,792 GB/s versus 273 GB/s).

The GB10-class hardware still wins for genuinely huge models or long-context workloads that don't fit on a single RTX 5090's 32GB, so the "best" hardware still depends entirely on what you're running.

Batched Throughput Changes Everything

Single-user chat and high-concurrency serving are different games entirely. The same RTX 5090 paired with vLLM can generate over 5,841 tokens per second at a batch size of 8, reportedly outperforming even data-center A100 GPUs by roughly 2.6x in that scenario. If you're serving multiple users, benchmark batched throughput specifically. Single-stream numbers won't tell you what you need to know.

Setting Up a Real Benchmark

Don't trust a single anecdotal number, including any of the ones above for your exact setup. Run your own test with your actual model and hardware.

Using llama.cpp Directly

./llama-bench -m your-model.gguf -p 512 -n 128

This measures prompt processing (-p, prefill speed) and generation (-n, decode speed) separately, since they behave differently and both matter for real use. If the flags look unfamiliar, our llama.cpp tutorial walks through the build first.

Using a Dedicated Benchmark Tool

Purpose-built CLI tools exist specifically for this now. One example runs structured, repeatable benchmarks against local or remote endpoints:

lmx benchmark run llama.cpp \
  --mode local \
  --model-path model.gguf \
  --prompt-tokens 512 \
  --output-tokens 128

Remote endpoint runs typically issue one untimed warmup request followed by several timed iterations, reporting median values plus per-run statistics like min, p50, mean, max, and standard deviation. That warmup step matters. Cold-start numbers include model loading time and misrepresent steady-state performance, so always discard the first run.

Using an Ollama-Based Dashboard

If you're already running Ollama, community tooling exists that wraps it in a monitoring dashboard, tracking time-to-first-token, tokens per second, tokens per minute, and live CPU/GPU/RAM usage during generation, then exports everything to CSV for comparison across models and hardware. This is the more visual option if you want to watch resource usage alongside throughput, rather than just reading final numbers.

What Actually Affects Your Numbers

Before you panic that your benchmark doesn't match published figures, check these variables:

  • Quantization level: Q4 versus Q8 versus full precision changes speed and memory dramatically; always compare like-for-like quant levels
  • Context length: longer context slows generation and eats more memory for the KV cache
  • Batch size: single-stream chat numbers differ wildly from batched serving numbers
  • Model architecture: a mixture-of-experts model often runs faster than its total parameter count suggests, since only a subset of experts activate per token
  • Warm vs cold model: first-run numbers include load time; always benchmark warm

FYI, if your numbers come in noticeably below published benchmarks on identical hardware, check your quantization format and context length first. Those two variables explain most mismatches. Our quantization explainer and GGUF format guide cover exactly what each level does to memory and speed.

Reading Coefficient of Variation

Good benchmarks report consistency, not just an average. One measured comparison found throughput coefficients of variation around 0.04% to 0.07% on one platform and 0.37% to 0.54% on another, meaning results were highly repeatable in both cases, just with genuinely different absolute speeds. A single benchmark run tells you almost nothing. Run at least three to five iterations and look at the spread before trusting any number, including your own.

A Practical Decision Framework

Decide your use case before shopping for hardware, not after:

  1. Casual single-user chat, 7B-8B models: an RTX 3060 12GB or an Apple M-series with 16-24GB unified memory handles this comfortably
  2. Larger models (32B+), single user: an RTX 4090 or an M4 Pro/Max with enough unified memory to hold the model
  3. Massive models (70B+) without multi-GPU: Apple Silicon with large unified memory wins here purely on capacity, even if per-token speed trails a top GPU
  4. Multi-user serving, high throughput: an RTX 5090 or better paired with vLLM for batched inference, not llama.cpp single-stream
  5. No hardware at all, want numbers first? Rent for an hour on a GPU rental service and benchmark before committing

IMO, decide your use case before shopping for hardware, not after. A setup optimized for one massive model running slowly looks completely different from one optimized for fast responses to many concurrent users.

Common Mistakes People Make

Comparing different quant levels

Recall the variables section directly — a Q2 number next to a Q4 number tells you nothing, and mixing them is how bad benchmarks get published.

Keeping the cold run

Recall the benchmark tooling section directly — the first attempt includes model load time, so it underreports steady-state speed every single time.

Trusting a single run

Recall the consistency section directly — without a spread across iterations, you can't tell a real difference from ordinary run-to-run noise.

Benchmarking chat instead of prefill and decode

Recall the llama.cpp section directly — prompt processing and generation behave differently, and a single blended number hides which one is actually your bottleneck.

Buying for peak specs

Recall the hardware tiers section directly — memory bandwidth and capacity relative to model size determine real speed long before the marketing TOPS number does.

Want to Go Deeper?

If you want structured practice on ML systems and evaluation, Educative's ML courses include hands-on labs that pair well with this kind of measurement work. The unlimited plan is useful when you're working through several model types in one stretch.

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is tokens per second in a local LLM?

Tokens per second is the rate at which a model generates output text after it begins responding. It is the number that determines whether chatting with a local model feels instant or painful, and it is measured separately from prompt processing speed.

What is time to first token?

Time to first token, or TTFT, is the delay between sending a prompt and the response starting to appear. A model with slow TTFT but fast generation feels sluggish to start and then snappy, so benchmark TTFT and generation speed together rather than either alone.

How do I benchmark my own local LLM?

Use llama-bench from llama.cpp with a prompt token count and a generation token count to measure prefill and decode separately, or a purpose-built CLI benchmark tool that runs warmup followed by several timed iterations and reports median values with spread statistics.

How many tokens per second do I need?

Roughly 10 tokens per second is usable for streaming chat, 20 to 30 feels comfortable, and beyond 50 the difference stops being noticeable while reading. Interactive use cares about time to first token more than raw peak throughput.

Why is my benchmark slower than published numbers?

Almost always quantization format and context length. Compare like-for-like quant levels, match context windows, confirm the model actually fits in VRAM rather than spilling to system memory, and discard the first cold run before comparing anything.

Does quantization affect speed?

Yes, substantially. Smaller quant levels use less memory bandwidth per token, which usually means faster generation, but quality drops. The right comparison is always between runs using the same quant level, since mixing them makes the numbers meaningless.

What is the difference between single-stream and batched throughput?

Single-stream measures one user generating text, which is what interactive chat depends on. Batched throughput measures many requests processed together and is what matters for serving multiple users. The same hardware can differ by an order of magnitude between the two.

Wrapping This Up

Tokens-per-second benchmarking cuts through vague marketing claims and tells you what a setup will actually feel like in daily use. Measure both time-to-first-token and steady-state tokens per second, always discard cold-start runs, match quantization levels when comparing hardware, and run multiple iterations to check consistency.

Will the exact numbers here match your specific setup? Probably not precisely — hardware, driver versions, and model builds all shift results. But run your own benchmark with a real tool tonight, using your actual model and quant level, and you'll know in minutes exactly what your hardware delivers instead of guessing from someone else's screenshot.

When the numbers point at a model you don't own yet, our local LLM tools comparison and LM Studio tutorial cover the runtimes worth measuring through.