[ ABORT TO HUD ]
SEQ. 1
SEQ. 2
Quantization Formats Compared
📐 Quantization Mastery⏱ 12 min⭐ 125 BASE XP
Why Quantize?
A 70B parameter model in FP16 requires ~140GB VRAM. Quantization reduces precision to fit models on smaller hardware while preserving quality.
The Decision Matrix
| Format | Best For | Key Advantage | Hardware |
|---|---|---|---|
| GGUF | Local / CPU / hybrid | Runs on anything (CPU, Mac, consumer GPU) | Universal |
| AWQ | Production GPU serving | Best quality at 4-bit, vLLM optimized | NVIDIA GPUs |
| GPTQ | Broad GPU inference | Wide ecosystem support, mature | NVIDIA GPUs |
| EXL2 | Maximum speed (single GPU) | Lowest latency for local high-end setups | High-end NVIDIA |
GGUF Quality Tiers
| Quant | Bits/Weight | Quality | 70B VRAM |
|---|---|---|---|
| Q8_0 | 8-bit | Near-lossless | ~70GB |
| Q6_K | 6-bit | Excellent | ~54GB |
| Q4_K_M | 4-bit | Great (recommended) | ~40GB |
| Q3_K_S | 3-bit | Acceptable | ~30GB |
| Q2_K | 2-bit | Quality cliff ⚠️ | ~20GB |
⚠️ The 4-Bit Rule: In 2026, 4-bit quantization is the industry standard. Going below 3-bit causes significant quality degradation (the "quality cliff"). If you have VRAM headroom, prefer Q6_K.
Calibration Best Practice
Post-training quantization quality depends on calibration data. For domain-specific use (medical, legal, coding), always calibrate with a sample of your actual production data rather than generic datasets.
KNOWLEDGE CHECK
QUERY 1 // 3
Which quantization format is recommended for production GPU serving with vLLM?
GGUF
AWQ
EXL2
FP16