Skip to content
fedi.software

Which open model fits your hardware — 2026

Open-weights models against popular graphics cards and computers: the recommended quantization (the largest up to Q8 that still gives 10+ tokens per second, otherwise a Q4-class one; below Q3 only when nothing else fits), where it runs (graphics memory, unified memory or offloaded to RAM) and an approximate generation speed.

Model Parameters
Qwen2.5 3B Instruct HF 3.1B Q8_0≈84 tok/s Fits in graphics memory Q8_0≈67 tok/s Fits in graphics memory Q8_0≈218 tok/s Fits in graphics memory Q8_0≈235 tok/s Fits in graphics memory Q8_0≈417 tok/s Fits in graphics memory Q8_0≈417 tok/s Fits in graphics memory Q8_0≈168 tok/s Fits in graphics memory Q8_0≈136 tok/s Fits in graphics memory Q8_0≈115 tok/s Fits in graphics memory Q8_0≈111 tok/s Fits in unified memory Q8_0≈167 tok/s Fits in unified memory Q8_0≈52 tok/s Fits in unified memory Q8_0≈56 tok/s Fits in unified memory Q8_0≈15.0 tok/s Fits in RAM (CPU) Q8_0≈34 tok/s Fits in RAM (CPU)
Llama 3.2 3B Instruct HF 3.2B Q8_0≈74 tok/s Fits in graphics memory Q8_0≈59 tok/s Fits in graphics memory Q8_0≈192 tok/s Fits in graphics memory Q8_0≈207 tok/s Fits in graphics memory Q8_0≈368 tok/s Fits in graphics memory Q8_0≈368 tok/s Fits in graphics memory Q8_0≈148 tok/s Fits in graphics memory Q8_0≈120 tok/s Fits in graphics memory Q8_0≈101 tok/s Fits in graphics memory Q8_0≈98 tok/s Fits in unified memory Q8_0≈147 tok/s Fits in unified memory Q8_0≈46 tok/s Fits in unified memory Q8_0≈49 tok/s Fits in unified memory Q8_0≈14.3 tok/s Fits in RAM (CPU) Q8_0≈32 tok/s Fits in RAM (CPU)
Gemma 3 4B HF 4.3B Q8_0≈67 tok/s Fits in graphics memory Q8_0≈54 tok/s Fits in graphics memory Q8_0≈175 tok/s Fits in graphics memory Q8_0≈189 tok/s Fits in graphics memory Q8_0≈335 tok/s Fits in graphics memory Q8_0≈335 tok/s Fits in graphics memory Q8_0≈135 tok/s Fits in graphics memory Q8_0≈109 tok/s Fits in graphics memory Q8_0≈92 tok/s Fits in graphics memory Q8_0≈89 tok/s Fits in unified memory Q8_0≈134 tok/s Fits in unified memory Q8_0≈42 tok/s Fits in unified memory Q8_0≈45 tok/s Fits in unified memory Q8_0≈13.8 tok/s Fits in RAM (CPU) Q8_0≈31 tok/s Fits in RAM (CPU)
Llama 3.1 8B Instruct HF 8B UD-Q6_K_XL≈37 tok/s Fits in graphics memory UD-Q6_K_XL≈29 tok/s Fits in graphics memory UD-Q6_K_XL≈95 tok/s Fits in graphics memory UD-Q6_K_XL≈103 tok/s Fits in graphics memory UD-Q6_K_XL≈182 tok/s Fits in graphics memory UD-Q6_K_XL≈182 tok/s Fits in graphics memory UD-Q6_K_XL≈73 tok/s Fits in graphics memory UD-Q6_K_XL≈59 tok/s Fits in graphics memory UD-Q6_K_XL≈50 tok/s Fits in graphics memory UD-Q6_K_XL≈49 tok/s Fits in unified memory UD-Q6_K_XL≈73 tok/s Fits in unified memory UD-Q6_K_XL≈23 tok/s Fits in unified memory UD-Q6_K_XL≈24 tok/s Fits in unified memory UD-Q6_K_XL≈10.3 tok/s Fits in RAM (CPU) UD-Q6_K_XL≈23 tok/s Fits in RAM (CPU)
Qwen3 8B HF 8.2B Q8_0≈31 tok/s Fits in graphics memory Q8_0≈25 tok/s Fits in graphics memory Q8_0≈80 tok/s Fits in graphics memory Q8_0≈87 tok/s Fits in graphics memory Q8_0≈154 tok/s Fits in graphics memory Q8_0≈154 tok/s Fits in graphics memory Q8_0≈62 tok/s Fits in graphics memory Q8_0≈50 tok/s Fits in graphics memory Q8_0≈42 tok/s Fits in graphics memory Q8_0≈41 tok/s Fits in unified memory Q8_0≈62 tok/s Fits in unified memory Q8_0≈19.2 tok/s Fits in unified memory Q8_0≈21 tok/s Fits in unified memory UD-Q6_K_XL≈10.2 tok/s Fits in RAM (CPU) Q8_0≈21 tok/s Fits in RAM (CPU)
Gemma 3 12B HF 12.2B UD-Q5_K_XL≈32 tok/s Fits in graphics memory Q8_0≈17.8 tok/s Fits in graphics memory Q8_0≈58 tok/s Fits in graphics memory Q8_0≈62 tok/s Fits in graphics memory Q8_0≈111 tok/s Fits in graphics memory Q8_0≈111 tok/s Fits in graphics memory Q8_0≈44 tok/s Fits in graphics memory Q8_0≈36 tok/s Fits in graphics memory Q8_0≈30 tok/s Fits in graphics memory Q8_0≈30 tok/s Fits in unified memory Q8_0≈44 tok/s Fits in unified memory Q8_0≈13.8 tok/s Fits in unified memory Q8_0≈14.8 tok/s Fits in unified memory Q4_1≈10.2 tok/s Fits in RAM (CPU) Q8_0≈17.1 tok/s Fits in RAM (CPU)
Qwen3 14B HF 14.8B Q4_K_M≈30 tok/s Fits in graphics memory Q6_K≈18.0 tok/s Fits in graphics memory Q8_0≈46 tok/s Fits in graphics memory Q8_0≈49 tok/s Fits in graphics memory Q8_0≈88 tok/s Fits in graphics memory Q8_0≈88 tok/s Fits in graphics memory Q8_0≈35 tok/s Fits in graphics memory Q8_0≈29 tok/s Fits in graphics memory Q8_0≈24 tok/s Fits in graphics memory Q8_0≈23 tok/s Fits in unified memory Q8_0≈35 tok/s Fits in unified memory Q8_0≈10.9 tok/s Fits in unified memory Q8_0≈11.7 tok/s Fits in unified memory UD-Q3_K_XL≈10.1 tok/s Fits in RAM (CPU) Q8_0≈14.6 tok/s Fits in RAM (CPU)
gpt-oss-20b HFMoE 20.9B / 3.6B MXFP4≈53 tok/s Runs with offload to RAM MXFP4≈55 tok/s Fits in graphics memory MXFP4162 tok/smeasured Fits in graphics memory MXFP4≈194 tok/s Fits in graphics memory MXFP4≈344 tok/s Fits in graphics memory MXFP4≈344 tok/s Fits in graphics memory MXFP4≈138 tok/s Fits in graphics memory MXFP4≈112 tok/s Fits in graphics memory MXFP4≈95 tok/s Fits in graphics memory MXFP4≈80 tok/s Fits in unified memory MXFP4116 tok/smeasured Fits in unified memory MXFP4≈59 tok/s Fits in unified memory MXFP4≈62 tok/s Fits in unified memory MXFP4≈17.2 tok/s Fits in RAM (CPU) MXFP4≈39 tok/s Fits in RAM (CPU)
Mistral Small 3.2 24B HF 24B Q3_K_M≈10.2 tok/s Runs with offload to RAM UD-Q3_K_XL≈18.4 tok/s Fits in graphics memory Q6_K≈37 tok/s Fits in graphics memory Q6_K≈40 tok/s Fits in graphics memory Q8_0≈56 tok/s Fits in graphics memory Q8_0≈56 tok/s Fits in graphics memory Q8_0≈22 tok/s Fits in graphics memory Q8_0≈18.2 tok/s Fits in graphics memory Q8_0≈15.3 tok/s Fits in graphics memory Q8_0≈14.9 tok/s Fits in unified memory Q8_0≈22 tok/s Fits in unified memory Q5_K_M≈10.3 tok/s Fits in unified memory Q5_K_M≈11.0 tok/s Fits in unified memory Q4_K_M≈6.9 tok/s Fits in RAM (CPU) Q8_0≈10.3 tok/s Fits in RAM (CPU)
Gemma 4 26B A4B HFMoE 25.8B / 3.8B UD-Q8_K_XL≈14.7 tok/s Runs with offload to RAM UD-Q3_K_M≈56 tok/s Fits in graphics memory UD-Q5_K_S≈129 tok/s Fits in graphics memory UD-Q5_K_S≈139 tok/s Fits in graphics memory UD-Q8_K_XL≈173 tok/s Fits in graphics memory UD-Q8_K_XL≈173 tok/s Fits in graphics memory UD-Q8_K_XL≈70 tok/s Fits in graphics memory UD-Q8_K_XL≈57 tok/s Fits in graphics memory UD-Q8_K_XL≈48 tok/s Fits in graphics memory UD-Q8_K_XL≈40 tok/s Fits in unified memory UD-Q8_K_XL≈60 tok/s Fits in unified memory UD-Q8_K_XL≈29 tok/s Fits in unified memory UD-Q8_K_XL≈31 tok/s Fits in unified memory UD-Q8_K_XL≈13.7 tok/s Fits in RAM (CPU) UD-Q8_K_XL≈31 tok/s Fits in RAM (CPU)
Gemma 3 27B HF 27.4B UD-IQ3_XXS≈12.7 tok/s Runs with offload to RAM Q3_K_S≈18.1 tok/s Fits in graphics memory UD-Q5_K_XL≈38 tok/s Fits in graphics memory UD-Q5_K_XL≈41 tok/s Fits in graphics memory UD-Q6_K_XL≈59 tok/s Fits in graphics memory Q8_0≈49 tok/s Fits in graphics memory Q8_0≈19.7 tok/s Fits in graphics memory Q8_0≈16.0 tok/s Fits in graphics memory Q8_0≈13.5 tok/s Fits in graphics memory Q8_0≈13.1 tok/s Fits in unified memory Q8_0≈19.6 tok/s Fits in unified memory Q4_1≈10.1 tok/s Fits in unified memory Q4_1≈10.8 tok/s Fits in unified memory Q4_K_M≈6.3 tok/s Fits in RAM (CPU) UD-Q6_K_XL≈10.8 tok/s Fits in RAM (CPU)
Qwen3 30B A3B HFMoE 30.5B / 3.3B Q8_0≈16.4 tok/s Runs with offload to RAM Q3_K_S≈66 tok/s Fits in graphics memory Q4_K_M154 tok/smeasured Fits in graphics memory Q5_K_S≈158 tok/s Fits in graphics memory UD-Q6_K_XL≈232 tok/s Fits in graphics memory Q8_0≈192 tok/s Fits in graphics memory Q8_0≈77 tok/s Fits in graphics memory Q8_0≈63 tok/s Fits in graphics memory Q8_0≈53 tok/s Fits in graphics memory Q8_0≈45 tok/s Fits in unified memory Q8_0≈67 tok/s Fits in unified memory Q8_0≈33 tok/s Fits in unified memory Q8_0≈35 tok/s Fits in unified memory Q8_0≈14.3 tok/s Fits in RAM (CPU) Q8_0≈32 tok/s Fits in RAM (CPU)
Gemma 4 31B HF 31.3B Q4_K_M≈3.8 tok/s Runs with offload to RAM Q3_K_S≈10.4 tok/s Runs with offload to RAM Q4_1≈37 tok/s Fits in graphics memory Q4_1≈40 tok/s Fits in graphics memory Q6_K≈55 tok/s Fits in graphics memory Q8_0≈43 tok/s Fits in graphics memory Q8_0≈17.1 tok/s Fits in graphics memory Q8_0≈13.9 tok/s Fits in graphics memory Q8_0≈11.7 tok/s Fits in graphics memory Q8_0≈11.3 tok/s Fits in unified memory Q8_0≈17.0 tok/s Fits in unified memory IQ4_XS≈10.3 tok/s Fits in unified memory Q4_K_S≈10.3 tok/s Fits in unified memory Q4_K_M≈5.7 tok/s Fits in RAM (CPU) Q6_K≈10.1 tok/s Fits in RAM (CPU)
Qwen3 32B HF 32.8B Q4_K_M≈3.7 tok/s Runs with offload to RAM UD-IQ3_XXS≈14.1 tok/s Runs with offload to RAM UD-Q4_K_XL≈35 tok/s Fits in graphics memory UD-Q4_K_XL≈38 tok/s Fits in graphics memory Q6_K≈51 tok/s Fits in graphics memory Q8_0≈40 tok/s Fits in graphics memory Q8_0≈16.0 tok/s Fits in graphics memory Q8_0≈13.0 tok/s Fits in graphics memory Q8_0≈11.0 tok/s Fits in graphics memory Q8_0≈10.6 tok/s Fits in unified memory Q8_0≈16.0 tok/s Fits in unified memory UD-Q3_K_XL≈10.3 tok/s Fits in unified memory IQ4_XS≈10.2 tok/s Fits in unified memory Q4_K_M≈5.4 tok/s Fits in RAM (CPU) Q5_K_M≈10.8 tok/s Fits in RAM (CPU)
Qwen3.5 35B A3B HFMoE 36B / 3B Q8_0≈17.9 tok/s Runs with offload to RAM Q8_0≈18.7 tok/s Runs with offload to RAM Q4_K_S≈191 tok/s Fits in graphics memory Q4_K_S≈205 tok/s Fits in graphics memory Q6_K≈274 tok/s Fits in graphics memory Q8_0≈220 tok/s Fits in graphics memory Q8_0≈89 tok/s Fits in graphics memory Q8_0≈72 tok/s Fits in graphics memory Q8_0≈61 tok/s Fits in graphics memory Q8_0≈51 tok/s Fits in unified memory Q8_0≈77 tok/s Fits in unified memory Q8_0≈37 tok/s Fits in unified memory Q8_0≈40 tok/s Fits in unified memory Q8_0≈15.0 tok/s Fits in RAM (CPU) Q8_0≈34 tok/s Fits in RAM (CPU)
Llama 3.3 70B Instruct HF 70.6B Q4_K_M≈1.3 tok/s Runs with offload to RAM Q4_K_M≈1.4 tok/s Runs with offload to RAM Q4_K_M≈2.0 tok/s Runs with offload to RAM Q4_K_M≈2.0 tok/s Runs with offload to RAM UD-IQ3_XXS≈49 tok/s Fits in graphics memory Q8_0≈18.8 tok/s Fits in graphics memory Q4_K_M16.3 tok/smeasured Fits in graphics memory Q4_1≈10.3 tok/s Fits in graphics memory IQ4_XS≈10.0 tok/s Fits in graphics memory UD-Q3_K_XL≈10.6 tok/s Fits in unified memory Q5_K_M≈11.2 tok/s Fits in unified memory Q4_K_M≈4.1 tok/s Fits in unified memory Q4_K_M≈4.4 tok/s Fits in unified memory Q4_K_M≈2.9 tok/s Fits in RAM (CPU) Q4_K_M≈6.6 tok/s Fits in RAM (CPU)
Qwen3 Next 80B A3B Instruct HFMoE 81.3B / 3B Q5_K_M≈24 tok/s Runs with offload to RAM Q6_K≈21 tok/s Runs with offload to RAM UD-Q6_K_XL≈26 tok/s Runs with offload to RAM UD-Q6_K_XL≈26 tok/s Runs with offload to RAM UD-Q6_K_XL≈31 tok/s Runs with offload to RAM Q8_0≈213 tok/s Fits in graphics memory Q4_K_S≈145 tok/s Fits in graphics memory UD-Q6_K_XL≈84 tok/s Fits in graphics memory Q8_0≈59 tok/s Fits in graphics memory Q8_0≈49 tok/s Fits in unified memory Q8_0≈74 tok/s Fits in unified memory Q8_0≈36 tok/s Fits in unified memory Q8_0≈39 tok/s Fits in unified memory Q8_0≈14.8 tok/s Fits in RAM (CPU) Q8_0≈33 tok/s Fits in RAM (CPU)
GLM 4.5 Air HFMoE 110B / 17B Q4_0≈5.3 tok/s Runs with offload to RAM Q4_K_S≈5.1 tok/s Runs with offload to RAM Q4_K_M≈5.6 tok/s Runs with offload to RAM Q4_K_M≈5.7 tok/s Runs with offload to RAM Q3_K_M≈10.2 tok/s Runs with offload to RAM Q5_K_M≈55 tok/s Fits in graphics memory Q4_0≈10.0 tok/s Runs with offload to RAM UD-Q4_K_XL≈22 tok/s Fits in graphics memory Q5_K_M≈15.2 tok/s Fits in graphics memory Q5_K_M≈12.8 tok/s Fits in unified memory Q8_0≈13.9 tok/s Fits in unified memory Q4_K_M≈10.6 tok/s Fits in unified memory Q5_K_M≈10.0 tok/s Fits in unified memory Q4_K_M≈8.0 tok/s Fits in RAM (CPU) Q8_0≈13.1 tok/s Fits in RAM (CPU)
gpt-oss-120b HFMoE 117B / 5.1B MXFP4≈19.1 tok/s Runs with offload to RAM MXFP4≈19.6 tok/s Runs with offload to RAM MXFP426–28 tok/smeasured Runs with offload to RAM MXFP4≈25 tok/s Runs with offload to RAM MXFP4≈31 tok/s Runs with offload to RAM MXFP4≈258 tok/s Fits in graphics memory MXFP4≈36 tok/s Runs with offload to RAM MXFP441–73 tok/smeasured Fits in graphics memory MXFP4≈71 tok/s Fits in graphics memory MXFP4≈60 tok/s Fits in unified memory MXFP4≈90 tok/s Fits in unified memory MXFP450 tok/smeasured Fits in unified memory MXFP461 tok/smeasured Fits in unified memory MXFP4≈15.8 tok/s Fits in RAM (CPU) MXFP4≈36 tok/s Fits in RAM (CPU)
Qwen3.5 122B A10B HFMoE 125B / 10B UD-IQ4_NL≈10.5 tok/s Runs with offload to RAM UD-IQ4_NL≈10.8 tok/s Runs with offload to RAM Q4_K_S≈11.1 tok/s Runs with offload to RAM Q4_K_S≈11.2 tok/s Runs with offload to RAM UD-Q4_K_XL≈11.9 tok/s Runs with offload to RAM UD-Q5_K_XL≈97 tok/s Fits in graphics memory UD-IQ3_XXS≈76 tok/s Fits in graphics memory UD-IQ4_NL≈46 tok/s Fits in graphics memory UD-Q5_K_XL≈27 tok/s Fits in graphics memory UD-Q5_K_XL≈23 tok/s Fits in unified memory Q8_0≈24 tok/s Fits in unified memory Q6_K≈15.1 tok/s Fits in unified memory Q6_K≈16.1 tok/s Fits in unified memory UD-Q5_K_XL≈10.4 tok/s Fits in RAM (CPU) Q8_0≈19.3 tok/s Fits in RAM (CPU)
Qwen3 235B A22B HFMoE 235B / 22B ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit UD-Q4_K_XL≈10.8 tok/s Runs with offload to RAM UD-Q2_K_XL≈8.0 tok/sstrong compression Runs with offload to RAM Q3_K_M≈6.1 tok/s Runs with offload to RAM UD-Q3_K_XL≈11.7 tok/s Runs with offload to RAM UD-Q2_K_XL≈19.4 tok/sstrong compression Fits in unified memory Q8_0≈10.8 tok/s Fits in unified memory UD-Q3_K_XL≈12.2 tok/s Fits in unified memory UD-Q3_K_XL≈13.0 tok/s Fits in unified memory Q4_K_M≈7.2 tok/s Fits in RAM (CPU) Q8_0≈10.9 tok/s Fits in RAM (CPU)
GLM 4.7 HFMoE 358B / 39.2B ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit UD-IQ3_XXS≈7.3 tok/s Runs with offload to RAM UD-IQ1_S≈5.6 tok/sstrong compression Runs with offload to RAM UD-IQ2_XXS≈4.7 tok/sstrong compression Runs with offload to RAM UD-IQ3_XXS≈3.5 tok/s Runs with offload to RAM UD-TQ1_0≈16.2 tok/sstrong compression Fits in unified memory Q4_1≈10.0 tok/s Fits in unified memory UD-IQ1_M≈9.6 tok/sstrong compression Fits in unified memory UD-IQ1_M≈10.3 tok/sstrong compression Fits in unified memory Q4_K_M≈4.7 tok/s Fits in RAM (CPU) Q4_1≈10.2 tok/s Fits in RAM (CPU)
DeepSeek R1 HFMoE 684B / 37B ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit UD-IQ1_S≈16.9 tok/sstrong compression Runs with offload to RAM ✕ Does not fit ✕ Does not fit UD-IQ1_S≈8.0 tok/sstrong compression Runs with offload to RAM ✕ Does not fit Q3_K_M≈14.9 tok/s Fits in unified memory ✕ Does not fit ✕ Does not fit Q4_K_M≈5.2 tok/s Fits in RAM (CPU) Q3_K_M≈13.9 tok/s Fits in RAM (CPU)
Kimi K2.5 HFMoE 1,027B / 32B ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit ✕ Does not fit UD-Q2_K_XL≈22 tok/sstrong compression Fits in unified memory ✕ Does not fit ✕ Does not fit Q3_K_S≈7.2 tok/s Fits in RAM (CPU) UD-IQ2_XXS≈19.7 tok/sstrong compression Fits in RAM (CPU)

What opens up at each memory size (context 8K)

  1. 8 GB Qwen2.5 3B Instruct · Llama 3.2 3B Instruct · Gemma 3 4B · Llama 3.1 8B Instruct · Qwen3 8B strong compression: Gemma 3 12B · Qwen3 14B
  2. 12 GB Gemma 3 12B · Qwen3 14B strong compression: Mistral Small 3.2 24B · Gemma 3 27B · Qwen3 30B A3B · Qwen3 32B
  3. 16 GB gpt-oss-20b · Mistral Small 3.2 24B · Gemma 4 26B A4B · Gemma 3 27B · Qwen3 30B A3B strong compression: Gemma 4 31B · Qwen3.5 35B A3B
  4. 24 GB Gemma 4 31B · Qwen3 32B · Qwen3.5 35B A3B strong compression: Llama 3.3 70B Instruct · Qwen3 Next 80B A3B Instruct
  5. 48 GB Llama 3.3 70B Instruct · Qwen3 Next 80B A3B Instruct · Qwen3.5 122B A10B strong compression: GLM 4.5 Air
  6. 96 GB GLM 4.5 Air · gpt-oss-120b strong compression: Qwen3 235B A22B · GLM 4.7
  7. 128 GB Qwen3 235B A22B
  8. 192 GB GLM 4.7 strong compression: DeepSeek R1
  9. 256 GB strong compression: Kimi K2.5
  10. 512 GB DeepSeek R1 · Kimi K2.5

≈ estimates for a 8K context: weights + KV cache + runtime overhead against the memory; speed from the memory bandwidth. Real numbers depend on the runtime and settings.

Updated · Sources: Hugging Face , www.nvidia.com , images.nvidia.com , support.apple.com , www.amd.com

Embed on your site

Paste this code where the infographic should appear: it updates by itself. Free to use — keep the attribution link.

JSON Our data: CC BY 4.0 with a link to the source; third-party data keeps the licence of its source.

What a cell tells you

Rows are well-known open-weights models, from about 3B to 1T parameters; columns are graphics cards, multi-GPU builds, Apple and AMD machines with unified memory, and CPU servers. Each cell names the quantization we recommend, its colour says where the model lives — graphics memory, unified memory, server RAM, offload to RAM, or does not fit — and the number is an approximate generation speed in tokens per second.

How the estimate is calculated

  1. Weights. Parameters times bits per weight, or the real size of the quant file from Hugging Face when it exists.
  2. KV cache. The memory that holds the conversation. The matrix assumes a context of 8192 tokens; at long contexts this part can outgrow a small model.
  3. Overhead. About 1.5 GB for the runtime and buffers, plus the vision projector of a VLM.
  4. Speed. Memory bandwidth divided by the bytes read for each token. A MoE model reads only its active parameters, which is why a large MoE can generate faster than a smaller dense model.

If the total exceeds the VRAM but fits together with system RAM, the cell shows offload: the model runs, only slower. The figures are estimates (≈), calibrated against measured runs and kept within about ±35% of them; a cell backed by a real measurement shows that value with a link to its source.

Why this quantization

We pick the largest quant up to Q8_0 that still reaches at least 10 tokens per second, otherwise a Q4-class file such as Q4_K_M. BF16 is never recommended — it doubles the memory for a gain you will hardly notice. Below Q3 we go only when nothing larger fits, and the cell then warns about strong compression.

Checking your own machine

The calculator under the matrix takes a model, a quant, a context length and your hardware, or simply your graphics memory, RAM and memory bandwidth. It draws the weights, KV cache and overhead as one bar against your capacity. The ladder below it lists what opens up at 8, 12, 16, 24, 48, 96 and 128 GB.

Two limits to keep in mind: the estimate covers generation, not the time spent reading a long prompt, and drivers, the runtime version and settings move real results either way. GGUF files run in llama.cpp, LM Studio, Ollama and Jan — see local AI apps. Models without safety filters are hidden behind a switch. For complete machines, go to the PC builds.