Which open model fits your hardware — 2026
Open-weights models against popular graphics cards and computers: the recommended quantization (the largest up to Q8 that still gives 10+ tokens per second, otherwise a Q4-class one; below Q3 only when nothing else fits), where it runs (graphics memory, unified memory or offloaded to RAM) and an approximate generation speed.
| Model | Parameters | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5 3B Instruct HF | 3.1B | Q8_0≈84 tok/s Fits in graphics memory | Q8_0≈67 tok/s Fits in graphics memory | Q8_0≈218 tok/s Fits in graphics memory | Q8_0≈235 tok/s Fits in graphics memory | Q8_0≈417 tok/s Fits in graphics memory | Q8_0≈417 tok/s Fits in graphics memory | Q8_0≈168 tok/s Fits in graphics memory | Q8_0≈136 tok/s Fits in graphics memory | Q8_0≈115 tok/s Fits in graphics memory | Q8_0≈111 tok/s Fits in unified memory | Q8_0≈167 tok/s Fits in unified memory | Q8_0≈52 tok/s Fits in unified memory | Q8_0≈56 tok/s Fits in unified memory | Q8_0≈15.0 tok/s Fits in RAM (CPU) | Q8_0≈34 tok/s Fits in RAM (CPU) |
| Llama 3.2 3B Instruct HF | 3.2B | Q8_0≈74 tok/s Fits in graphics memory | Q8_0≈59 tok/s Fits in graphics memory | Q8_0≈192 tok/s Fits in graphics memory | Q8_0≈207 tok/s Fits in graphics memory | Q8_0≈368 tok/s Fits in graphics memory | Q8_0≈368 tok/s Fits in graphics memory | Q8_0≈148 tok/s Fits in graphics memory | Q8_0≈120 tok/s Fits in graphics memory | Q8_0≈101 tok/s Fits in graphics memory | Q8_0≈98 tok/s Fits in unified memory | Q8_0≈147 tok/s Fits in unified memory | Q8_0≈46 tok/s Fits in unified memory | Q8_0≈49 tok/s Fits in unified memory | Q8_0≈14.3 tok/s Fits in RAM (CPU) | Q8_0≈32 tok/s Fits in RAM (CPU) |
| Gemma 3 4B HF | 4.3B | Q8_0≈67 tok/s Fits in graphics memory | Q8_0≈54 tok/s Fits in graphics memory | Q8_0≈175 tok/s Fits in graphics memory | Q8_0≈189 tok/s Fits in graphics memory | Q8_0≈335 tok/s Fits in graphics memory | Q8_0≈335 tok/s Fits in graphics memory | Q8_0≈135 tok/s Fits in graphics memory | Q8_0≈109 tok/s Fits in graphics memory | Q8_0≈92 tok/s Fits in graphics memory | Q8_0≈89 tok/s Fits in unified memory | Q8_0≈134 tok/s Fits in unified memory | Q8_0≈42 tok/s Fits in unified memory | Q8_0≈45 tok/s Fits in unified memory | Q8_0≈13.8 tok/s Fits in RAM (CPU) | Q8_0≈31 tok/s Fits in RAM (CPU) |
| Llama 3.1 8B Instruct HF | 8B | UD-Q6_K_XL≈37 tok/s Fits in graphics memory | UD-Q6_K_XL≈29 tok/s Fits in graphics memory | UD-Q6_K_XL≈95 tok/s Fits in graphics memory | UD-Q6_K_XL≈103 tok/s Fits in graphics memory | UD-Q6_K_XL≈182 tok/s Fits in graphics memory | UD-Q6_K_XL≈182 tok/s Fits in graphics memory | UD-Q6_K_XL≈73 tok/s Fits in graphics memory | UD-Q6_K_XL≈59 tok/s Fits in graphics memory | UD-Q6_K_XL≈50 tok/s Fits in graphics memory | UD-Q6_K_XL≈49 tok/s Fits in unified memory | UD-Q6_K_XL≈73 tok/s Fits in unified memory | UD-Q6_K_XL≈23 tok/s Fits in unified memory | UD-Q6_K_XL≈24 tok/s Fits in unified memory | UD-Q6_K_XL≈10.3 tok/s Fits in RAM (CPU) | UD-Q6_K_XL≈23 tok/s Fits in RAM (CPU) |
| Qwen3 8B HF | 8.2B | Q8_0≈31 tok/s Fits in graphics memory | Q8_0≈25 tok/s Fits in graphics memory | Q8_0≈80 tok/s Fits in graphics memory | Q8_0≈87 tok/s Fits in graphics memory | Q8_0≈154 tok/s Fits in graphics memory | Q8_0≈154 tok/s Fits in graphics memory | Q8_0≈62 tok/s Fits in graphics memory | Q8_0≈50 tok/s Fits in graphics memory | Q8_0≈42 tok/s Fits in graphics memory | Q8_0≈41 tok/s Fits in unified memory | Q8_0≈62 tok/s Fits in unified memory | Q8_0≈19.2 tok/s Fits in unified memory | Q8_0≈21 tok/s Fits in unified memory | UD-Q6_K_XL≈10.2 tok/s Fits in RAM (CPU) | Q8_0≈21 tok/s Fits in RAM (CPU) |
| Gemma 3 12B HF | 12.2B | UD-Q5_K_XL≈32 tok/s Fits in graphics memory | Q8_0≈17.8 tok/s Fits in graphics memory | Q8_0≈58 tok/s Fits in graphics memory | Q8_0≈62 tok/s Fits in graphics memory | Q8_0≈111 tok/s Fits in graphics memory | Q8_0≈111 tok/s Fits in graphics memory | Q8_0≈44 tok/s Fits in graphics memory | Q8_0≈36 tok/s Fits in graphics memory | Q8_0≈30 tok/s Fits in graphics memory | Q8_0≈30 tok/s Fits in unified memory | Q8_0≈44 tok/s Fits in unified memory | Q8_0≈13.8 tok/s Fits in unified memory | Q8_0≈14.8 tok/s Fits in unified memory | Q4_1≈10.2 tok/s Fits in RAM (CPU) | Q8_0≈17.1 tok/s Fits in RAM (CPU) |
| Qwen3 14B HF | 14.8B | Q4_K_M≈30 tok/s Fits in graphics memory | Q6_K≈18.0 tok/s Fits in graphics memory | Q8_0≈46 tok/s Fits in graphics memory | Q8_0≈49 tok/s Fits in graphics memory | Q8_0≈88 tok/s Fits in graphics memory | Q8_0≈88 tok/s Fits in graphics memory | Q8_0≈35 tok/s Fits in graphics memory | Q8_0≈29 tok/s Fits in graphics memory | Q8_0≈24 tok/s Fits in graphics memory | Q8_0≈23 tok/s Fits in unified memory | Q8_0≈35 tok/s Fits in unified memory | Q8_0≈10.9 tok/s Fits in unified memory | Q8_0≈11.7 tok/s Fits in unified memory | UD-Q3_K_XL≈10.1 tok/s Fits in RAM (CPU) | Q8_0≈14.6 tok/s Fits in RAM (CPU) |
| gpt-oss-20b HFMoE | 20.9B / 3.6B | MXFP4≈53 tok/s Runs with offload to RAM | MXFP4≈55 tok/s Fits in graphics memory | MXFP4162 tok/smeasured Fits in graphics memory | MXFP4≈194 tok/s Fits in graphics memory | MXFP4≈344 tok/s Fits in graphics memory | MXFP4≈344 tok/s Fits in graphics memory | MXFP4≈138 tok/s Fits in graphics memory | MXFP4≈112 tok/s Fits in graphics memory | MXFP4≈95 tok/s Fits in graphics memory | MXFP4≈80 tok/s Fits in unified memory | MXFP4116 tok/smeasured Fits in unified memory | MXFP4≈59 tok/s Fits in unified memory | MXFP4≈62 tok/s Fits in unified memory | MXFP4≈17.2 tok/s Fits in RAM (CPU) | MXFP4≈39 tok/s Fits in RAM (CPU) |
| Mistral Small 3.2 24B HF | 24B | Q3_K_M≈10.2 tok/s Runs with offload to RAM | UD-Q3_K_XL≈18.4 tok/s Fits in graphics memory | Q6_K≈37 tok/s Fits in graphics memory | Q6_K≈40 tok/s Fits in graphics memory | Q8_0≈56 tok/s Fits in graphics memory | Q8_0≈56 tok/s Fits in graphics memory | Q8_0≈22 tok/s Fits in graphics memory | Q8_0≈18.2 tok/s Fits in graphics memory | Q8_0≈15.3 tok/s Fits in graphics memory | Q8_0≈14.9 tok/s Fits in unified memory | Q8_0≈22 tok/s Fits in unified memory | Q5_K_M≈10.3 tok/s Fits in unified memory | Q5_K_M≈11.0 tok/s Fits in unified memory | Q4_K_M≈6.9 tok/s Fits in RAM (CPU) | Q8_0≈10.3 tok/s Fits in RAM (CPU) |
| Gemma 4 26B A4B HFMoE | 25.8B / 3.8B | UD-Q8_K_XL≈14.7 tok/s Runs with offload to RAM | UD-Q3_K_M≈56 tok/s Fits in graphics memory | UD-Q5_K_S≈129 tok/s Fits in graphics memory | UD-Q5_K_S≈139 tok/s Fits in graphics memory | UD-Q8_K_XL≈173 tok/s Fits in graphics memory | UD-Q8_K_XL≈173 tok/s Fits in graphics memory | UD-Q8_K_XL≈70 tok/s Fits in graphics memory | UD-Q8_K_XL≈57 tok/s Fits in graphics memory | UD-Q8_K_XL≈48 tok/s Fits in graphics memory | UD-Q8_K_XL≈40 tok/s Fits in unified memory | UD-Q8_K_XL≈60 tok/s Fits in unified memory | UD-Q8_K_XL≈29 tok/s Fits in unified memory | UD-Q8_K_XL≈31 tok/s Fits in unified memory | UD-Q8_K_XL≈13.7 tok/s Fits in RAM (CPU) | UD-Q8_K_XL≈31 tok/s Fits in RAM (CPU) |
| Gemma 3 27B HF | 27.4B | UD-IQ3_XXS≈12.7 tok/s Runs with offload to RAM | Q3_K_S≈18.1 tok/s Fits in graphics memory | UD-Q5_K_XL≈38 tok/s Fits in graphics memory | UD-Q5_K_XL≈41 tok/s Fits in graphics memory | UD-Q6_K_XL≈59 tok/s Fits in graphics memory | Q8_0≈49 tok/s Fits in graphics memory | Q8_0≈19.7 tok/s Fits in graphics memory | Q8_0≈16.0 tok/s Fits in graphics memory | Q8_0≈13.5 tok/s Fits in graphics memory | Q8_0≈13.1 tok/s Fits in unified memory | Q8_0≈19.6 tok/s Fits in unified memory | Q4_1≈10.1 tok/s Fits in unified memory | Q4_1≈10.8 tok/s Fits in unified memory | Q4_K_M≈6.3 tok/s Fits in RAM (CPU) | UD-Q6_K_XL≈10.8 tok/s Fits in RAM (CPU) |
| Qwen3 30B A3B HFMoE | 30.5B / 3.3B | Q8_0≈16.4 tok/s Runs with offload to RAM | Q3_K_S≈66 tok/s Fits in graphics memory | Q4_K_M154 tok/smeasured Fits in graphics memory | Q5_K_S≈158 tok/s Fits in graphics memory | UD-Q6_K_XL≈232 tok/s Fits in graphics memory | Q8_0≈192 tok/s Fits in graphics memory | Q8_0≈77 tok/s Fits in graphics memory | Q8_0≈63 tok/s Fits in graphics memory | Q8_0≈53 tok/s Fits in graphics memory | Q8_0≈45 tok/s Fits in unified memory | Q8_0≈67 tok/s Fits in unified memory | Q8_0≈33 tok/s Fits in unified memory | Q8_0≈35 tok/s Fits in unified memory | Q8_0≈14.3 tok/s Fits in RAM (CPU) | Q8_0≈32 tok/s Fits in RAM (CPU) |
| Gemma 4 31B HF | 31.3B | Q4_K_M≈3.8 tok/s Runs with offload to RAM | Q3_K_S≈10.4 tok/s Runs with offload to RAM | Q4_1≈37 tok/s Fits in graphics memory | Q4_1≈40 tok/s Fits in graphics memory | Q6_K≈55 tok/s Fits in graphics memory | Q8_0≈43 tok/s Fits in graphics memory | Q8_0≈17.1 tok/s Fits in graphics memory | Q8_0≈13.9 tok/s Fits in graphics memory | Q8_0≈11.7 tok/s Fits in graphics memory | Q8_0≈11.3 tok/s Fits in unified memory | Q8_0≈17.0 tok/s Fits in unified memory | IQ4_XS≈10.3 tok/s Fits in unified memory | Q4_K_S≈10.3 tok/s Fits in unified memory | Q4_K_M≈5.7 tok/s Fits in RAM (CPU) | Q6_K≈10.1 tok/s Fits in RAM (CPU) |
| Qwen3 32B HF | 32.8B | Q4_K_M≈3.7 tok/s Runs with offload to RAM | UD-IQ3_XXS≈14.1 tok/s Runs with offload to RAM | UD-Q4_K_XL≈35 tok/s Fits in graphics memory | UD-Q4_K_XL≈38 tok/s Fits in graphics memory | Q6_K≈51 tok/s Fits in graphics memory | Q8_0≈40 tok/s Fits in graphics memory | Q8_0≈16.0 tok/s Fits in graphics memory | Q8_0≈13.0 tok/s Fits in graphics memory | Q8_0≈11.0 tok/s Fits in graphics memory | Q8_0≈10.6 tok/s Fits in unified memory | Q8_0≈16.0 tok/s Fits in unified memory | UD-Q3_K_XL≈10.3 tok/s Fits in unified memory | IQ4_XS≈10.2 tok/s Fits in unified memory | Q4_K_M≈5.4 tok/s Fits in RAM (CPU) | Q5_K_M≈10.8 tok/s Fits in RAM (CPU) |
| Qwen3.5 35B A3B HFMoE | 36B / 3B | Q8_0≈17.9 tok/s Runs with offload to RAM | Q8_0≈18.7 tok/s Runs with offload to RAM | Q4_K_S≈191 tok/s Fits in graphics memory | Q4_K_S≈205 tok/s Fits in graphics memory | Q6_K≈274 tok/s Fits in graphics memory | Q8_0≈220 tok/s Fits in graphics memory | Q8_0≈89 tok/s Fits in graphics memory | Q8_0≈72 tok/s Fits in graphics memory | Q8_0≈61 tok/s Fits in graphics memory | Q8_0≈51 tok/s Fits in unified memory | Q8_0≈77 tok/s Fits in unified memory | Q8_0≈37 tok/s Fits in unified memory | Q8_0≈40 tok/s Fits in unified memory | Q8_0≈15.0 tok/s Fits in RAM (CPU) | Q8_0≈34 tok/s Fits in RAM (CPU) |
| Llama 3.3 70B Instruct HF | 70.6B | Q4_K_M≈1.3 tok/s Runs with offload to RAM | Q4_K_M≈1.4 tok/s Runs with offload to RAM | Q4_K_M≈2.0 tok/s Runs with offload to RAM | Q4_K_M≈2.0 tok/s Runs with offload to RAM | UD-IQ3_XXS≈49 tok/s Fits in graphics memory | Q8_0≈18.8 tok/s Fits in graphics memory | Q4_K_M16.3 tok/smeasured Fits in graphics memory | Q4_1≈10.3 tok/s Fits in graphics memory | IQ4_XS≈10.0 tok/s Fits in graphics memory | UD-Q3_K_XL≈10.6 tok/s Fits in unified memory | Q5_K_M≈11.2 tok/s Fits in unified memory | Q4_K_M≈4.1 tok/s Fits in unified memory | Q4_K_M≈4.4 tok/s Fits in unified memory | Q4_K_M≈2.9 tok/s Fits in RAM (CPU) | Q4_K_M≈6.6 tok/s Fits in RAM (CPU) |
| Qwen3 Next 80B A3B Instruct HFMoE | 81.3B / 3B | Q5_K_M≈24 tok/s Runs with offload to RAM | Q6_K≈21 tok/s Runs with offload to RAM | UD-Q6_K_XL≈26 tok/s Runs with offload to RAM | UD-Q6_K_XL≈26 tok/s Runs with offload to RAM | UD-Q6_K_XL≈31 tok/s Runs with offload to RAM | Q8_0≈213 tok/s Fits in graphics memory | Q4_K_S≈145 tok/s Fits in graphics memory | UD-Q6_K_XL≈84 tok/s Fits in graphics memory | Q8_0≈59 tok/s Fits in graphics memory | Q8_0≈49 tok/s Fits in unified memory | Q8_0≈74 tok/s Fits in unified memory | Q8_0≈36 tok/s Fits in unified memory | Q8_0≈39 tok/s Fits in unified memory | Q8_0≈14.8 tok/s Fits in RAM (CPU) | Q8_0≈33 tok/s Fits in RAM (CPU) |
| GLM 4.5 Air HFMoE | 110B / 17B | Q4_0≈5.3 tok/s Runs with offload to RAM | Q4_K_S≈5.1 tok/s Runs with offload to RAM | Q4_K_M≈5.6 tok/s Runs with offload to RAM | Q4_K_M≈5.7 tok/s Runs with offload to RAM | Q3_K_M≈10.2 tok/s Runs with offload to RAM | Q5_K_M≈55 tok/s Fits in graphics memory | Q4_0≈10.0 tok/s Runs with offload to RAM | UD-Q4_K_XL≈22 tok/s Fits in graphics memory | Q5_K_M≈15.2 tok/s Fits in graphics memory | Q5_K_M≈12.8 tok/s Fits in unified memory | Q8_0≈13.9 tok/s Fits in unified memory | Q4_K_M≈10.6 tok/s Fits in unified memory | Q5_K_M≈10.0 tok/s Fits in unified memory | Q4_K_M≈8.0 tok/s Fits in RAM (CPU) | Q8_0≈13.1 tok/s Fits in RAM (CPU) |
| gpt-oss-120b HFMoE | 117B / 5.1B | MXFP4≈19.1 tok/s Runs with offload to RAM | MXFP4≈19.6 tok/s Runs with offload to RAM | MXFP426–28 tok/smeasured Runs with offload to RAM | MXFP4≈25 tok/s Runs with offload to RAM | MXFP4≈31 tok/s Runs with offload to RAM | MXFP4≈258 tok/s Fits in graphics memory | MXFP4≈36 tok/s Runs with offload to RAM | MXFP441–73 tok/smeasured Fits in graphics memory | MXFP4≈71 tok/s Fits in graphics memory | MXFP4≈60 tok/s Fits in unified memory | MXFP4≈90 tok/s Fits in unified memory | MXFP450 tok/smeasured Fits in unified memory | MXFP461 tok/smeasured Fits in unified memory | MXFP4≈15.8 tok/s Fits in RAM (CPU) | MXFP4≈36 tok/s Fits in RAM (CPU) |
| Qwen3.5 122B A10B HFMoE | 125B / 10B | UD-IQ4_NL≈10.5 tok/s Runs with offload to RAM | UD-IQ4_NL≈10.8 tok/s Runs with offload to RAM | Q4_K_S≈11.1 tok/s Runs with offload to RAM | Q4_K_S≈11.2 tok/s Runs with offload to RAM | UD-Q4_K_XL≈11.9 tok/s Runs with offload to RAM | UD-Q5_K_XL≈97 tok/s Fits in graphics memory | UD-IQ3_XXS≈76 tok/s Fits in graphics memory | UD-IQ4_NL≈46 tok/s Fits in graphics memory | UD-Q5_K_XL≈27 tok/s Fits in graphics memory | UD-Q5_K_XL≈23 tok/s Fits in unified memory | Q8_0≈24 tok/s Fits in unified memory | Q6_K≈15.1 tok/s Fits in unified memory | Q6_K≈16.1 tok/s Fits in unified memory | UD-Q5_K_XL≈10.4 tok/s Fits in RAM (CPU) | Q8_0≈19.3 tok/s Fits in RAM (CPU) |
| Qwen3 235B A22B HFMoE | 235B / 22B | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | UD-Q4_K_XL≈10.8 tok/s Runs with offload to RAM | UD-Q2_K_XL≈8.0 tok/sstrong compression Runs with offload to RAM | Q3_K_M≈6.1 tok/s Runs with offload to RAM | UD-Q3_K_XL≈11.7 tok/s Runs with offload to RAM | UD-Q2_K_XL≈19.4 tok/sstrong compression Fits in unified memory | Q8_0≈10.8 tok/s Fits in unified memory | UD-Q3_K_XL≈12.2 tok/s Fits in unified memory | UD-Q3_K_XL≈13.0 tok/s Fits in unified memory | Q4_K_M≈7.2 tok/s Fits in RAM (CPU) | Q8_0≈10.9 tok/s Fits in RAM (CPU) |
| GLM 4.7 HFMoE | 358B / 39.2B | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | UD-IQ3_XXS≈7.3 tok/s Runs with offload to RAM | UD-IQ1_S≈5.6 tok/sstrong compression Runs with offload to RAM | UD-IQ2_XXS≈4.7 tok/sstrong compression Runs with offload to RAM | UD-IQ3_XXS≈3.5 tok/s Runs with offload to RAM | UD-TQ1_0≈16.2 tok/sstrong compression Fits in unified memory | Q4_1≈10.0 tok/s Fits in unified memory | UD-IQ1_M≈9.6 tok/sstrong compression Fits in unified memory | UD-IQ1_M≈10.3 tok/sstrong compression Fits in unified memory | Q4_K_M≈4.7 tok/s Fits in RAM (CPU) | Q4_1≈10.2 tok/s Fits in RAM (CPU) |
| DeepSeek R1 HFMoE | 684B / 37B | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | UD-IQ1_S≈16.9 tok/sstrong compression Runs with offload to RAM | ✕ Does not fit | ✕ Does not fit | UD-IQ1_S≈8.0 tok/sstrong compression Runs with offload to RAM | ✕ Does not fit | Q3_K_M≈14.9 tok/s Fits in unified memory | ✕ Does not fit | ✕ Does not fit | Q4_K_M≈5.2 tok/s Fits in RAM (CPU) | Q3_K_M≈13.9 tok/s Fits in RAM (CPU) |
| Kimi K2.5 HFMoE | 1,027B / 32B | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | ✕ Does not fit | UD-Q2_K_XL≈22 tok/sstrong compression Fits in unified memory | ✕ Does not fit | ✕ Does not fit | Q3_K_S≈7.2 tok/s Fits in RAM (CPU) | UD-IQ2_XXS≈19.7 tok/sstrong compression Fits in RAM (CPU) |
Calculator: will it fit?
What opens up at each memory size (context 8K)
- 8 GB Qwen2.5 3B Instruct · Llama 3.2 3B Instruct · Gemma 3 4B · Llama 3.1 8B Instruct · Qwen3 8B strong compression: Gemma 3 12B · Qwen3 14B
- 12 GB Gemma 3 12B · Qwen3 14B strong compression: Mistral Small 3.2 24B · Gemma 3 27B · Qwen3 30B A3B · Qwen3 32B
- 16 GB gpt-oss-20b · Mistral Small 3.2 24B · Gemma 4 26B A4B · Gemma 3 27B · Qwen3 30B A3B strong compression: Gemma 4 31B · Qwen3.5 35B A3B
- 24 GB Gemma 4 31B · Qwen3 32B · Qwen3.5 35B A3B strong compression: Llama 3.3 70B Instruct · Qwen3 Next 80B A3B Instruct
- 48 GB Llama 3.3 70B Instruct · Qwen3 Next 80B A3B Instruct · Qwen3.5 122B A10B strong compression: GLM 4.5 Air
- 96 GB GLM 4.5 Air · gpt-oss-120b strong compression: Qwen3 235B A22B · GLM 4.7
- 128 GB Qwen3 235B A22B
- 192 GB GLM 4.7 strong compression: DeepSeek R1
- 256 GB strong compression: Kimi K2.5
- 512 GB DeepSeek R1 · Kimi K2.5
≈ estimates for a 8K context: weights + KV cache + runtime overhead against the memory; speed from the memory bandwidth. Real numbers depend on the runtime and settings.
Updated · Sources: Hugging Face , www.nvidia.com , images.nvidia.com , support.apple.com , www.amd.com
Embed on your site
Paste this code where the infographic should appear: it updates by itself. Free to use — keep the attribution link.
What a cell tells you
Rows are well-known open-weights models, from about 3B to 1T parameters; columns are graphics cards, multi-GPU builds, Apple and AMD machines with unified memory, and CPU servers. Each cell names the quantization we recommend, its colour says where the model lives — graphics memory, unified memory, server RAM, offload to RAM, or does not fit — and the number is an approximate generation speed in tokens per second.
How the estimate is calculated
- Weights. Parameters times bits per weight, or the real size of the quant file from Hugging Face when it exists.
- KV cache. The memory that holds the conversation. The matrix assumes a context of 8192 tokens; at long contexts this part can outgrow a small model.
- Overhead. About 1.5 GB for the runtime and buffers, plus the vision projector of a VLM.
- Speed. Memory bandwidth divided by the bytes read for each token. A MoE model reads only its active parameters, which is why a large MoE can generate faster than a smaller dense model.
If the total exceeds the VRAM but fits together with system RAM, the cell shows offload: the model runs, only slower. The figures are estimates (≈), calibrated against measured runs and kept within about ±35% of them; a cell backed by a real measurement shows that value with a link to its source.
Why this quantization
We pick the largest quant up to Q8_0 that still reaches at least 10 tokens per second, otherwise a Q4-class file such as Q4_K_M. BF16 is never recommended — it doubles the memory for a gain you will hardly notice. Below Q3 we go only when nothing larger fits, and the cell then warns about strong compression.
Checking your own machine
The calculator under the matrix takes a model, a quant, a context length and your hardware, or simply your graphics memory, RAM and memory bandwidth. It draws the weights, KV cache and overhead as one bar against your capacity. The ladder below it lists what opens up at 8, 12, 16, 24, 48, 96 and 128 GB.
Two limits to keep in mind: the estimate covers generation, not the time spent reading a long prompt, and drivers, the runtime version and settings move real results either way. GGUF files run in llama.cpp, LM Studio, Ollama and Jan — see local AI apps. Models without safety filters are hidden behind a switch. For complete machines, go to the PC builds.