Contents
Figure 1: The machine on your desk is the whole data center
Buying a laptop for local LLM work is genuinely different from buying one for gaming or video editing, and most spec sheets won't tell you what actually matters. The number you should care about isn't the GPU name or the marketing TOPS figure. It's memory bandwidth, and once you understand why, the whole buying decision gets a lot clearer.
I'll walk through the real trade-offs, with actual measured numbers from recent 2026 benchmarks, not vendor marketing.
The One Concept That Decides Everything
Local LLM performance is memory-bandwidth-bound, not compute-bound. Generating each token means reading the entire model's weights from memory, so the speed at which your hardware can move data, not how many operations per second it can theoretically do, sets your ceiling.
This explains a genuinely surprising result several reviewers hit in 2026: Apple Silicon can outrun a high-end RTX 5090 on the biggest local models, purely because of memory architecture. One hands-on comparison found that once you step up to a large model like Qwen3-Coder-Next at FP8, taking up 85GB, the RTX 5090 isn't even in the same conversation, since its 32GB of VRAM simply can't hold the model at all, while a Mac with unified memory can.
The flip side matters too: on models that fit comfortably in 32GB, the RTX 5090's raw bandwidth advantage (roughly 1,792 GB/s versus an M5 Max's 614 GB/s) means it decodes noticeably faster. Neither platform wins universally. It depends entirely on whether your target model fits in the smaller, faster pool or needs the bigger, somewhat slower one.
If you've never measured your own hardware, our local LLM benchmarking guide shows how to get real tokens-per-second numbers before you spend anything.
Reading the Memory Bandwidth Numbers
Here's the landscape as of late 2026, compiled from several independent benchmark sources. Treat exact figures as approximate since methodology (quantization level, specific model, warm vs cold) varies between them:
| Platform | Memory Bandwidth | Practical Ceiling |
|---|---|---|
| RTX 5090 (32GB GDDR7) | ~1,792 GB/s | Fastest for models that fit under 32GB |
| RTX 5090 Mobile | ~896–1,150 GB/s | Strong laptop option, still VRAM-limited |
| Apple M5 Max (up to 128GB) | ~614 GB/s | Best combination of speed and capacity for 70B+ models |
| Apple M4 Max (up to 128GB) | ~546 GB/s | Still excellent, one generation back |
| Apple M5 Pro (up to 48GB) | ~273 GB/s | Good for 13B–32B class models |
| x86 LPDDR5X (integrated) | ~120 GB/s | The bottleneck on most non-Apple thin-and-lights |
That last row matters a lot if you're shopping outside Apple's lineup or a discrete-GPU gaming laptop. A thin, non-gaming Windows laptop with only integrated graphics will bottleneck hard, since standard LPDDR5X bandwidth is roughly 5 to 15x slower than either a discrete GPU or Apple's unified memory.
What Actually Fits in Memory
Before comparing speed, check whether a model fits at all. Rough memory requirements at common Q4 quantization:
- 7B–8B models: 6–8 GB
- 13B–14B models: 10–12 GB
- 32B models: 20–24 GB
- 70B models: 40–48 GB (won't fit on a 32GB RTX 5090 at all; needs a 64GB+ unified-memory Mac, or heavier quantization)
- 120B+ MoE models: 80GB+, realistically only 96–128GB unified memory Macs or multi-GPU servers
That 70B row is the one that trips people up. A 70B model at Q4 needs roughly 35GB, which exceeds a 32GB RTX 5090's VRAM entirely, forcing you either into aggressive Q3 quantization with real quality loss, or onto a Mac with enough unified memory to hold it comfortably.
Want to check what your own machine is actually handing over to a model? Two commands tell you most of what you need:
# Ollama: live memory and throughput for loaded models
ollama ps
# Windows: total system RAM in bytes
powershell -c "(Get-CimInstance Win32_ComputerSystem).TotalPhysicalMemory"
# macOS
sysctl -n hw.memsize
And to measure the machine you already own before deciding whether to upgrade:
llama-bench -m model-q4_k_m.gguf -p 512 -n 128
The Two Real Paths: Apple Silicon or Discrete NVIDIA
Apple Silicon: MacBook Pro (M4 Max / M5 Max)
Best for developers who want maximum model capacity, silent operation, and genuinely good battery life while running local models. A MacBook Pro with an M4 Max or M5 Max and 64GB–128GB of unified memory can hold 70B and even some 120B-class MoE models entirely in memory, something no laptop-class discrete GPU can currently match, and it does this while running silently and unplugged.
Real measured numbers vary source to source, but land roughly here: an M4 Max with 64GB runs Llama 3 70B at Q4 around 18 to 78 tokens per second depending on the specific benchmark and model variant, and the newer M5 generation pushes further, with every GPU core gaining a dedicated Neural Accelerator and base memory bandwidth rising to around 153 GB/s on the Air tier alone.
The trade-off: software maturity. NVIDIA still dominates for training and fine-tuning, with CUDA, PyTorch, vLLM, and TensorRT-LLM all built around it first. Apple's MLX ecosystem is genuinely good and growing fast, but it's not the CUDA equivalent in scope yet, and plenty of local LLM tooling still expects Metal directly rather than a higher-level abstraction.
Where to buy (affiliate links):
- MacBook Pro M4 Max with 64GB+ unified memory — the 70B-on-the-go pick, silent and unplugged
- MacBook Pro M5 Max with up to 128GB — newest generation, dedicated Neural Accelerator per GPU core
- MacBook Pro M4 Pro with 32GB — the mid-range developer pick for 13B–32B models
- MacBook Air M5 with 32GB — light 7B–14B work with the best battery life
- MacBook Air with 24GB — the cheapest comfortable way into local inference
Discrete NVIDIA: RTX 5090 Laptops
Best for developers who want raw speed on models that fit, plus the ability to actually train or fine-tune something afterward without switching machines. An RTX 5090 mobile setup can push memory bandwidth over 1,150 GB/s, translating to genuinely excellent tokens-per-second on quantized models up to the low-30B range.
The trade-off: capacity and power. A 70B model simply won't fit in 32GB of VRAM without heavy quantization, and running near peak performance means staying plugged in, with audible fans. One head-to-head test found a Razer Blade 16 with an RTX 5090 hit around 110 tok/s on a 70B-class model via CUDA, but only while plugged in with fans running, compared to a MacBook Pro M4 Max running the same model class silently on battery. Power draw differs by roughly 10x per session between the two approaches.
A Third Path: AMD Unified Memory (Ryzen AI Max)
Worth knowing about if you're not locked into either ecosystem: AMD's Ryzen AI Max+ 395 platform, found in machines like the HP ZBook Ultra, brings a unified-memory approach similar in spirit to Apple's, with configurations reaching 128GB. One 2026 benchmark ranked it as the top pick specifically because that memory capacity let it handle larger models than most discrete-GPU Windows laptops at a lower price point than a comparable Mac.
It's a genuinely interesting middle path if you want Windows/Linux flexibility without Apple's software ecosystem, though it's newer and less battle-tested than either Apple Silicon or NVIDIA.
My Recommendation by Use Case
- Mobile developer, mostly 7B–14B models, wants battery life: A MacBook Pro with M4 Pro/M5 Pro and 32GB unified memory, or a MacBook Air M5 with 32GB for lighter work
- Need to run 70B+ class models on the go, silently: MacBook Pro M4 Max or MacBook Pro M5 Max with 64GB–128GB unified memory
- Want maximum raw speed on models under 32GB, plus fine-tuning capability: A laptop with an RTX 5090 mobile GPU and 24GB+ system RAM
- Budget-conscious, Windows/Linux preference, want capacity over raw speed: An AMD Ryzen AI Max+ machine with 64GB–128GB unified memory
- Just experimenting, not doing serious edge AI dev: A Mac Mini or MacBook Air with 24–32GB handles 7B–14B models comfortably and costs far less than the flagship options above
IMO, if you don't already know you need CUDA specifically (for training, not just inference), the unified-memory path, Apple or AMD, gives you more model capacity per dollar and a genuinely nicer day-to-day experience for local LLM work. Reach for NVIDIA specifically when fine-tuning or training is a real part of your workflow, not just running inference.
A Practical Decision Framework
Let me save you some research time with straightforward guidance:
- Know your target model size first — buy for the largest model you actually intend to run, not the one you might someday try
- Under 32B, need speed and CUDA? RTX 5090 laptop; it's the fastest path for models that fit
- 70B or larger, want silence and battery? MacBook Pro Max with 64GB+ unified memory
- On a budget, need Windows/Linux and capacity? AMD Ryzen AI Max+ with 64GB+ unified memory
- Only running 7B–14B? Don't buy a flagship at all — 32GB on a Mac Mini or mini PC is plenty
- Training or fine-tuning in scope? CUDA still wins; check our best GPUs for deep learning guide for the desktop equivalent
The mistake isn't choosing the wrong brand — it's buying a fast GPU that can't hold the model you actually want to run.
Common Mistakes People Make
Buying for the GPU name, not the memory
Recall the bandwidth section — a marketing TOPS figure says nothing about how fast you'll actually decode tokens.
Undersizing memory "to save money"
Recall the fit table — the gap between "comfortable" and "can't load the model" is a hard wall, not a gradual slowdown.
Ignoring thermals
Recall the NVIDIA section — long generation sessions push laptops for minutes at a time, and throttling can cut tokens/sec by 30-50% on inadequately cooled machines.
Falling for NPU marketing
Recall the buying tips — NPU acceleration of larger LLMs is still mostly a 2026–2027 software story; GPU and unified memory remain what actually matters today.
Budgeting around the wrong quantization
Recall the fit table — Q4_K_M is the community standard for a reason: it's the sweet spot between quality and size, so do your memory math around it.
Trusting MSRP
Recall the current market — 2026 has seen real DRAM supply pressure pushing some flagship GPU laptop prices up meaningfully, so check current street prices rather than MSRP.
Recommended Books
- Deep Learning by Ian Goodfellow, Yoshua Bengio, and Aaron Courville — the theory behind why every generated token has to re-read the weights, which is the whole reason this buying guide exists.
- Computer Organization and Design: RISC-V Edition by David A. Patterson and John L. Hennessy — memory hierarchies explained properly, so bandwidth numbers on a spec sheet finally mean something.
- CUDA by Example by Jason Sanders and Edward Kandrot — if you go the NVIDIA route, this is the practical companion for understanding what CUDA actually buys you beyond inference.
Want to Go Deeper?
If you want structured practice on ML systems and deployment, Educative's ML courses include hands-on labs that pair well with hardware decisions like these. The unlimited plan is useful when you're working through several targets in one stretch.
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What matters most when buying a laptop for local LLM work?
Memory bandwidth first, memory capacity second, and only then the GPU name. Generating each token means reading the whole model from memory, so bandwidth sets your tokens-per-second ceiling and capacity decides whether the model loads at all.
How much memory do I need for a 70B model?
Roughly 40-48GB at Q4 quantization. That exceeds a 32GB RTX 5090 entirely, so a 70B model needs a 64GB or larger unified-memory Mac, an AMD Ryzen AI Max machine with enough capacity, or heavy Q3 quantization with real quality loss.
Is Apple Silicon or NVIDIA better for running local LLMs?
It depends on your target model size. Apple Silicon wins on capacity and efficiency: it can hold 70B and some 120B-class models that no laptop GPU fits, running silently on battery. NVIDIA wins on raw speed for models that fit in 32GB and on training and fine-tuning, where CUDA still dominates.
Are gaming laptops good for local LLM development?
They can be, if the discrete GPU has enough VRAM for your models, since bandwidth and CUDA support are strong. The trade-offs are capacity, power draw, and noise: 70B models will not fit in 32GB, peak speed usually requires staying plugged in, and fans run under long generation sessions.
Do NPU TOPS figures matter when choosing a laptop?
Not yet for larger local LLMs. NPU acceleration of substantial models is still mostly a 2026-2027 software story, while GPU compute and memory bandwidth are what determine real tokens-per-second today. Ignore TOPS marketing for now and buy memory instead.
How much RAM do I need for 7B to 14B models?
Between 6GB and 12GB of free memory at Q4 quantization, so 32GB of system memory gives you comfortable headroom for the model plus context and your other work. That class of model runs well on a MacBook Air or a modest Windows laptop.
Is AMD Ryzen AI Max a good option for local LLMs?
It is the most interesting middle path. Machines like the HP ZBook Ultra with Ryzen AI Max+ 395 reach up to 128GB of unified memory, handling larger models than most discrete-GPU Windows laptops at a lower price than a comparable Mac, though the ecosystem is newer and less battle-tested.
Wrapping This Up
The best laptop for local LLM work isn't the one with the flashiest spec sheet, it's the one whose memory bandwidth and capacity actually match the models you plan to run. Apple Silicon wins on capacity and efficiency for big models, discrete NVIDIA wins on raw speed for models that fit and on training flexibility, and unified-memory Windows alternatives are carving out a real middle ground.
Will one platform be objectively "best" forever? No, this space moves fast, and by the time you're reading this, another generation may have shifted the numbers again. But understanding memory bandwidth as the deciding factor, rather than chasing a GPU name on a spec sheet, will serve you well no matter which generation you're shopping in.
Once you've picked the machine, our best hardware for edge AI and TinyML and best GPUs for running local LLMs at home guides cover the desktop side of the same decision.