Sam Austin AI

Best GPUs for Running Local LLMs at Home

September 7, 2026 9 min read Sam Austin
Contents
NVIDIA GPU comparison for running local LLMs with Ollama and llama.cpp at home
NVIDIA GPU comparison for running local LLMs with Ollama and llama.cpp at home

Here's the anecdote that should genuinely settle every GPU debate before it starts: someone bought a fast card with too little memory, watched a 32B model crawl, returned it, and bought a used RTX 3090 instead. Same model, cheaper card, entirely in VRAM — and it ran at usable speed. That's the whole guide in one story. Compute is the last thing worth worrying about here.

This is the concrete GPU-buying companion to the mini PC and hardware guides from earlier in this series, narrowed specifically to the question everyone in the local LLM arc of this series eventually asks: which actual graphics card should I buy? Recall the mini PC article's finding that RAM matters more than CPU for that category — the GPU version of that lesson is even starker, and the numbers behind it are genuinely dramatic.

By the end of this guide, you'll know exactly how much VRAM your target model needs, which specific card fits your budget, and why the cliff you fall off when you get this wrong is so much steeper than people expect. IMO, this is one of those rare hardware categories where the buying rule really is this simple :)

The One Rule That Actually Matters: VRAM Capacity First

If your model doesn't fit in VRAM, performance doesn't degrade gracefully — it collapses. This is the single most important thing to internalize before looking at any spec sheet.

  • An RTX 5090 running Llama 3.3 70B fully in VRAM hits 45+ tokens per second. The same card, same model, offloading to system RAM instead: 1-2 tokens per second — slower than reading speed.
  • This isn't a gradual slowdown, it's a cliff — 5 to 20 times slower once a model starts spilling into system RAM, not a modest percentage hit.
  • GPU memory needs to cover more than just the model weights — working memory, KV cache, and serving overhead all compete for that same VRAM budget, which is why you generally want headroom beyond the bare minimum a weights-only calculation suggests.

The practical rule of thumb, repeated consistently across every current source on this topic: decide the model size you want to run, read the required VRAM off a chart, then buy a card one tier up for headroom. Compute speed genuinely comes second.

How Much VRAM Different Model Sizes Actually Need

This connects directly to the GGUF and quantization articles from earlier in this series — these numbers assume Q4-class quantization, the genuine sweet spot covered there.

  • 7B models — comfortably fit in 8-12GB, genuinely modest requirements.
  • 13B-14B models — want roughly 12-16GB for comfortable headroom.
  • 32B-34B models (Qwen 32B, DeepSeek-R1 32B) — need roughly 19GB at Q4 — this is exactly why a 24GB card is the community's repeatedly-cited sweet spot for this specific tier.
  • 70B models — need roughly 40GB at 4-bit for weights alone, before context and concurrency — this genuinely rules out any single 24GB consumer card.
  • 70B at full FP16 — needs roughly 140GB — squarely data-center territory, not a home setup consideration at all.

2026's goalpost has genuinely moved from a couple years ago — 8GB was a reasonable entry point in 2024; 16GB is now basically the practical minimum for a usable local LLM experience, driven by how much LLM capability has expanded even at the same parameter counts.

The Value Pick: Used RTX 3090 (24GB)

This is genuinely the most consistently recommended card across every current source for anyone serious about running real models at home without spending flagship money.

The RTX 3090 delivers 24GB of VRAM — enough to comfortably handle the entire 32B-class tier at Q4, which represents genuinely the best quality you can run comfortably on one affordable card.

  • Roughly $900 used, making this the pragmatic path to high-parameter capability without paying flagship pricing.
  • The "NVIDIA Tax" avoidance angle matters here too — 24GB for hundreds less than a 4090 or 5090.
  • This is the card the returned-fast-card story above ends with — genuinely representative of what actually happens when people prioritize VRAM over raw speed after learning the lesson the hard way.

The Budget Entry Points

If your ambitions are genuinely modest — 7B models, basic experimentation — you don't need to spend 3090 money at all.

  • RTX 4070 Ti (~$600) — delivers roughly 80 tokens/sec on 7B models, about 80% of a 4090's performance at roughly a third of the cost. The clear pick if you're confident you'll stay in the 7B range.
  • RTX 5060 Ti 16GB — widely cited as the current value entry point into local AI, genuinely capable of running the smaller end of the model spectrum comfortably.
  • Intel Arc B580 (12GB) — the budget entry for basic 7B models specifically, worth knowing about if you're not committed to the NVIDIA/CUDA ecosystem and want the cheapest functional starting point.

Most people starting fresh with local AI should genuinely buy a value-tier card and spend the savings on fast system RAM and storage rather than stretching for a flagship card upfront — a recommendation that shows up consistently across current buying guides.

The High-End Consumer Tier: RTX 4090 and RTX 5090

Once your ambitions genuinely exceed 32B-class models, the consumer ceiling gets more expensive fast.

  • RTX 4090 (24GB, ~$1,800) — delivers roughly 150 tokens/sec on 7B models, and handles the same 32-34B tier as the 3090 but meaningfully faster — the pick if speed genuinely matters and 24GB remains your VRAM ceiling.
  • RTX 5090 (32GB GDDR7, ~current flagship pricing) — the only consumer GeForce card that lets 30B-class models run entirely in VRAM and lets 70B-class models run with only minimal offloading, something no other single consumer card manages. Fastest gaming GPU on the market as a side benefit, with 21,760 CUDA cores.
  • Two RTX 4090s pooled (48GB combined, ~$3,600 total) — runs Llama 3.3 70B at Q5 quantization around 100 tokens/sec — genuinely the best consumer multi-GPU option for reaching 70B territory without moving to workstation-tier hardware.

Avoid attempting 70B on a single consumer GPU regardless of card — the Q2 quantization required to force-fit it degrades quality severely, well past the point where the exercise is worthwhile.

The Workstation and Prosumer Tier

For genuinely serious 70B+ work or small-team always-on serving, a different tier exists above consumer GeForce cards entirely.

  • RTX PRO 6000 (workstation-class, larger VRAM ceiling) — described directly by Linus Tech Tips' February 2026 hands-on: it "can fit much larger models in VRAM than the 5090 could ever dream of running efficiently" — genuinely the card's entire value proposition in one sentence. At roughly 94% of the 5090's raw throughput but a substantially larger VRAM ceiling, it enables 70B at Q6 or Q8 (meaningfully better output quality for coding and reasoning tasks) or 120B-class models at Q4, plus genuine multi-user serving from a single card.
  • ECC memory matters at this tier for long-running workloads where bit-flip errors would otherwise compound over extended sessions — a genuinely different reliability consideration than anything in the consumer tier.
  • A100 or similar data-center cards — the recommended tier for small-team, always-on serving where consumer cards' reliability and memory ceiling genuinely become limiting.
  • H100, H200, B200 — production/enterprise scale for concurrent multi-user throughput, bought as servers rather than desktop components entirely.

The AMD Question

Worth addressing directly since it comes up constantly: AMD is genuinely viable for local LLMs in 2026, with real caveats worth understanding upfront.

  • The RX 7900 XTX (24GB) competes well on price, fitting the same dense 32-34B tier at Q5 as an RTX 4090, plus handling aggressive quantizations of newer architectures like Llama 4 Scout.
  • The tradeoff is genuinely software, not hardware — most LLM frameworks (vLLM, llama.cpp, Ollama, exactly the tools covered earlier in this series) optimize for CUDA first. ROCm support requires more driver and framework configuration effort, and inference throughput on ROCm typically trails CUDA by 10-20% at the same VRAM tier.
  • Check the current llama.cpp/Ollama/vLLM ROCm compatibility list for your specific card before buying — this genuinely isn't a universal guarantee across AMD's entire lineup, echoing the same ROCm caveat flagged in this series' Ollama and Stable Diffusion tutorials.

Quick Decision Table

Your GoalRecommended CardVRAM
7B models only, tightest budgetRTX 4070 Ti or RTX 5060 Ti12-16GB
Best value for 32B-class daily driverUsed RTX 309024GB
32B-class, maximum speedRTX 409024GB
70B with minimal offloading, single cardRTX 509032GB
70B fully in VRAM, best qualityTwo RTX 4090s pooled, or RTX PRO 600048GB+
Small-team always-on servingRTX PRO 6000 or A10048-96GB
Production/enterprise concurrent scaleH100 / H200 / B200 (server deployment)80GB+

Apple Silicon: A Genuinely Different Path Worth Knowing

Recall the local LLM tools article's note that Ollama's Mac backend runs on MLX rather than llama.cpp specifically because unified memory exploits differently than a discrete GPU. This matters again here: a Mac with sufficient unified memory (128GB configurations exist) offers a genuinely different route to running large models — no discrete GPU VRAM ceiling in the traditional sense, at the cost of the raw throughput a dedicated NVIDIA card delivers. Worth a dedicated comparison if you're choosing between a Mac and a PC build specifically for this purpose rather than assuming one is universally better.

Common Mistakes People Make

  • Buying based on compute benchmarks (CUDA cores, TFLOPS) before checking VRAM capacity. The cliff from fitting-in-VRAM to spilling-into-RAM dwarfs any compute difference between cards — VRAM capacity is the first filter, not the last.
  • Attempting 70B on a single 24GB consumer card. The resulting Q2 quantization degrades quality severely — either accept a smaller model or budget for a multi-GPU or workstation-tier solution.
  • Assuming AMD's ROCm compatibility matches CUDA's maturity across every framework. Check the current compatibility list for llama.cpp, Ollama, and vLLM specifically before committing to an AMD card.
  • Overbuying flagship compute for a 7B-only use case. A budget-tier card like the RTX 4070 Ti delivers 80% of flagship performance at a third of the cost for models in this size range — match spending to your actual model ceiling.
  • Ignoring that GPU memory needs headroom beyond bare model weights. KV cache, context length, and serving overhead all compete for the same VRAM — buy one tier up from a weights-only calculation, not exactly at the calculated minimum.

Wrapping This Up

Choosing a GPU for local LLMs in 2026 comes down to a genuinely simple decision process, however much marketing noise surrounds it: decide the model size you actually want to run, check its VRAM requirement at Q4 quantization, and buy a card with that much capacity plus headroom — everything else is secondary. The used RTX 3090 remains the community's consistent value pick for the 32B sweet spot, while the RTX 5090 or a multi-GPU setup is genuinely necessary once 70B becomes the target.

Remember that the VRAM cliff is dramatic, not gradual — 5 to 20x slower the moment a model spills into system RAM — and that compute specs are worth comparing only after VRAM capacity has already ruled cards in or out. FYI, this closes out the local LLM hardware thread running through this whole series — from Ollama's software layer through GGUF's quantization format to this article's actual "which card do I put in my computer" answer :)

Now go check the exact VRAM requirement for the specific model you actually want to run — not a generic "7B vs 70B" category, but the real Q4_K_M file size from Hugging Face — before you buy anything. That one number, more than any benchmark chart in this article, is what should decide your purchase.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles