Sam Austin AI

Model Quantization Explained: Shrink Neural Networks Without Losing Accuracy (2026)

September 7, 2026 14 min read Sam Austin
Contents
Model Quantization Explained Shrink Neural Networks INT8 INT4 PTQ QAT
Model Quantization Explained Shrink Neural Networks INT8 INT4 PTQ QAT

Figure 1: Model quantization — compressing neural networks without sacrificing practical accuracy

Here's the number that should genuinely surprise you: a 70B-parameter model in fp16 needs 140GB of memory and won't fit on a single H100. The same model in int4 needs just 35GB and runs comfortably on that same GPU — with 3-4x faster decode throughput on top. Same weights, same architecture, same capability in practice. Just... smaller numbers representing them.

This concept has been lurking underneath nearly every article in this series without getting its own proper treatment — it's the actual mechanism behind llama.cpp's Q4_K_M tag, the "reduced precision" step in the Arduino TinyML conversion, and the Edge TPU-specific model files from the Raspberry Pi tutorial. This is genuinely the article that connects all of those dots, explaining the one technique doing more work than any other single idea in this whole edge AI and local LLM series.

By the end of this guide, you'll understand exactly what quantization does mathematically, the two major approaches to doing it, and why the accuracy cost isn't linear with how aggressively you compress. IMO, this is one of those ideas that seems like it shouldn't work at all until you actually see the numbers :)

What Quantization Actually Does

Quantization compresses a trained neural network's weights — and sometimes its activations too — from a high-precision numeric format down to a lower-precision one. Typically that means going from 32-bit or 16-bit floating point (FP32, FP16, BF16) down to 8-bit or 4-bit integers (INT8, INT4), or newer low-precision float formats (FP8, FP4).

The point isn't really disk size. For large models, it's the memory footprint and the bandwidth needed to actually feed those weights into the processor doing the math — smaller numbers move faster through memory.

INT8 quantization typically costs well under 1% accuracy on most models — genuinely close to free, given the size and speed benefits gained.

INT4 with a good method costs roughly 1-3% on standard benchmarks, and has become the production sweet spot for memory-bound serving in 2026.

This is the exact same principle behind every "compressed model" moment in this series — llama.cpp's Q4_K_M tag, the Arduino tutorial's Optimize.DEFAULT conversion flag, the Raspberry Pi Edge TPU's specifically-compiled model files. Different hardware, different scale, identical underlying idea.

The Two Fundamental Approaches: PTQ vs. QAT

Every quantization technique falls into one of two camps, and understanding the distinction explains almost every tradeoff you'll encounter.

Post-Training Quantization (PTQ)

PTQ takes an already-trained, full-precision model and converts it to lower precision afterward — no retraining involved at all.

The process typically computes the actual range of values in the model's weights using a small calibration dataset, then determines a scale factor mapping that range onto the target lower-precision format.

This is genuinely the low-effort path — no access to the original training pipeline or dataset required, just the trained model and a modest calibration set.

The tradeoff: PTQ can cause meaningful performance degradation specifically for models that are highly sensitive to precision loss — some architectures tolerate this conversion gracefully, others genuinely don't.

Quantization-Aware Training (QAT)

QAT bakes the quantization process directly into training itself, rather than applying it as an afterthought.

It works by inserting "fake" or simulated quantization operations into the model's computation graph during training — these take full-precision values, simulate the rounding and clamping effects of converting to something like INT8, then convert back to full precision for the next layer.

This forces the training algorithm to learn weights that are genuinely resilient to quantization's information loss — the model adapts around the constraint instead of just enduring it after the fact.

QAT is more complex and requires a full retraining cycle, but it almost always yields higher accuracy than PTQ, often approaching the original full-precision model's performance closely.

The practical rule of thumb: reach for PTQ first since it's dramatically cheaper to apply. When PTQ results in unacceptable quality loss, QAT is the actual fix — not a fancier PTQ calibration technique, but a genuinely different, more expensive process.

Why Quality Loss Isn't Linear With Bit-Width

Here's the part that trips up intuition: you might expect going from 8 bits to 4 bits to cost roughly twice the accuracy loss of going from 16 to 8. It genuinely doesn't work that way.

INT8 is essentially free on most models — sub-1% perplexity hit, a rounding error in practical terms.

INT4 with a good PTQ method costs roughly 1-3% — still a fully reasonable production tradeoff for most use cases.

Drop to 3 bits, and the loss curve bends sharply. Outlier channels that were perfectly tolerable at 4 bits become genuinely catastrophic at 3 — this is exactly where you typically need QAT or mixed-precision tricks rather than a simple PTQ pass.

This nonlinearity is why "just quantize more aggressively" isn't a free lever you can pull indefinitely. There's a genuine cliff, not a smooth slope, and it shows up at a different bit-width depending on the specific model architecture.

The Outlier Problem, and How Modern Methods Solve It

The reason naive quantization sometimes fails badly comes down to a specific, well-understood problem: a small fraction of weight channels matter disproportionately more than the rest, and naive uniform scaling handles them badly.

AWQ (Activation-aware Weight Quantization) identifies which weight channels matter most and protects them specifically with per-channel scaling, rather than treating every channel identically.

SmoothQuant takes a different angle on the same problem — it migrates quantization difficulty from activations to weights through per-channel rescaling, enabling full INT8 weights-and-activations quantization without needing QAT at all.

GPTQ is another widely used PTQ method specifically for large language models, using a more sophisticated optimization process than simple range-based calibration to minimize the resulting quantization error.

These methods are all attacking the exact same underlying problem — outlier channels that wreck naive scaling — from genuinely different angles. None of them is universally "best"; which one wins depends on your specific model architecture and target bit-width.

Weights vs. Activations: A Distinction Worth Understanding

Quantization discussions often mention "weights" and "activations" separately, and the distinction genuinely matters for what you actually gain.

Weight quantization saves memory footprint on the device — the model file itself is smaller, and less data needs to move through memory bandwidth during inference.

Activation quantization improves computational efficiency further by enabling integer multiplication in the actual compute, rather than just shrinking storage.

"W8A8" or "W8A16" notation you'll see in technical documentation specifies exactly this — 8-bit weights with 8-bit or 16-bit activations respectively, letting you mix precision levels for different parts of the same pipeline.

Mixed precision is genuinely common in practice — some layers quantize cleanly to INT4, while more sensitive layers stay at INT8 or even full precision, blending aggressiveness with accuracy protection exactly where each layer needs it.

Real Numbers: What Quantization Actually Buys You

Beyond the general accuracy figures, concrete measured results help ground how significant this technique actually is in practice.

Quantization has demonstrated up to 68% reduction in model size while maintaining performance within single-digit percentage points of full-precision baselines, using proper scaling techniques.

INT8 quantization has shown roughly 40% reduction in computational cost and power consumption, with INT4 pushing that further to around 60%.

On edge devices specifically, INT8 configurations have delivered roughly 2.4x throughput improvement, with INT4 reaching around 3x — genuinely substantial gains for hardware with real power and thermal constraints.

On data-center GPUs like the RTX 5090, dedicated Tensor Cores optimized specifically for INT8 matrix multiplication deliver a step-function throughput improvement over FP16 — this isn't a marginal software optimization, it's hardware built specifically to exploit lower precision.

Quantization Across the Hardware Spectrum

Here's where this concept threads through nearly everything covered earlier in this series, at genuinely different scales.

TinyML on Arduino used quantization to fit a model into kilobytes of tensor arena — the most extreme end of this spectrum, where every byte saved matters enormously.

The Raspberry Pi Coral accelerator specifically required Edge TPU-compiled model variants — a fixed-function ASIC that only accelerates a specific quantization scheme, INT8 in that case.

llama.cpp's GGUF quantization levels (Q4_K_M, Q6_K, Q8_0) are exactly this same technique applied to large language models — the naming convention just makes the specific bit allocation and method explicit in the filename itself.

Large-scale LLM serving on data-center GPUs uses the exact same INT4/INT8 principles discussed here, just at a scale where 140GB versus 35GB determines whether a model fits on one GPU or needs several.

The underlying math is identical at every one of these scales — only the specific tools, target hardware, and acceptable accuracy tradeoffs change.

A Practical Decision Framework

If you're deciding how to quantize your own model, here's a sensible sequence to follow.

Start with PTQ at INT8. This is close to free in accuracy terms for most architectures, and requires no retraining — always the first thing to try.

If you need smaller/faster and INT8 alone isn't enough, try INT4 with a modern PTQ method like GPTQ or AWQ rather than naive uniform quantization — these specifically handle the outlier-channel problem that naive approaches struggle with.

If INT4 PTQ produces unacceptable quality loss, move to QAT rather than pushing PTQ further — this is genuinely the point where the more expensive retraining-based approach becomes worth its cost.

Consider mixed precision (keeping sensitive layers at higher precision while aggressively quantizing the rest) before assuming your entire model needs uniform treatment.

Always validate on your actual task and data, not just a generic benchmark — the accuracy cliff's exact location genuinely varies by architecture and use case.

Where Quantization Fits in the Bigger Picture

If you've been following this series, quantization is the technique connecting the Arduino TinyML tutorial's Optimize.DEFAULT flag, the Raspberry Pi ML deploy guide's Edge TPU model files, and the llama.cpp tutorial's GGUF quantization levels. Same math, different scales — from kilobytes on a microcontroller to terabytes across data-center clusters.

For a deeper dive into running quantized models locally, our Ollama setup guide walks through pulling pre-quantized models with one command. The local LLM tools comparison covers how Ollama, LM Studio, and Jan each handle quantization selection differently under the hood.

Next Steps: Master Model Optimization

Ready to optimize models for production deployment? Educative offers interactive courses on model optimization, inference acceleration, and building efficient ML systems — practice in real sandboxed environments and learn by building actual optimization pipelines.

Common Mistakes People Make

Assuming accuracy loss scales linearly with bit-width reduction. The real curve has a genuine cliff, typically somewhere around 3 bits for many architectures — don't extrapolate linearly from the INT8-to-INT4 cost.

Reaching for QAT by default instead of trying PTQ first. QAT's retraining cost is real; PTQ solves the large majority of quantization needs for a fraction of the effort.

Ignoring the outlier channel problem. Naive uniform quantization without AWQ, SmoothQuant, or similar techniques can produce far worse results than the bit-width alone would suggest.

Quantizing without validating on your actual downstream task. A generic perplexity or benchmark number doesn't guarantee your specific use case tolerates the same accuracy hit.

Forgetting that hardware needs to actually support your chosen precision. A model quantized to a format your target chip's Tensor Cores or NPU doesn't accelerate natively gains far less than the theoretical numbers suggest.

Wrapping This Up

Quantization is genuinely the single technique doing the most invisible heavy lifting across this entire edge AI and local LLM series — compressing weights (and often activations) from high-precision floats down to compact integers, trading a small, carefully managed amount of accuracy for dramatic gains in size, speed, and power efficiency. PTQ handles the vast majority of real-world cases cheaply; QAT exists for the harder cases where that quality loss genuinely isn't acceptable.

Remember that the accuracy cost isn't linear — INT8 is nearly free, INT4 is a very reasonable tradeoff, and going lower without techniques like AWQ or QAT risks a genuine accuracy cliff rather than a gentle decline. FYI, the next time you pick a Q4_K_M tag in llama.cpp or see an Optimize.DEFAULT flag in a TensorFlow Lite conversion, you now know exactly what's happening underneath that one line :)

Now go back to the llama.cpp tutorial from earlier in this series and actually compare Q8_0 against Q4_K_M on the same model using llama-bench. Seeing the real size-versus-quality tradeoff on your own hardware is worth more than any benchmark table in this article.

Share this article X Facebook LinkedIn Reddit WhatsApp

Related Articles