Skip to content
fedi.software

PC builds for local AI models, 20B to 120B — 2026

From one 16 GB graphics card to three RTX 3090s, an EPYC server and unified-memory machines: what each build runs and how fast, with the source of every measurement.

  • One 16 GB GPU

    Budget
    Budget
    Graphics memory
    ≈16 GB
    RAM
    ≈32 GB
    Components
    • GPU: RTX 5060 Ti 16 GB / RTX 4060 Ti 16 GB
    • CPU: any 6–8-core desktop CPU
    • Motherboard: any ATX/mATX with a PCIe x16 slot
    • RAM: 32 GB DDR4/DDR5
    • Power supply: 550–650 W
    • PCIe: 1× x16 (x8 electrical is enough)
    Runs
    Notes
    gpt-oss-20b (≈12 GB) fits entirely in VRAM with room for a long context. Qwen3-30B-A3B Q4_K_M (≈18.6 GB) does not fit: a few expert layers go to system RAM with --n-cpu-moe, still usable because only ~3B parameters are active. Measured on an RTX 5060 Ti 16 GB (llama-bench tg128).
  • One RTX 3090 24 GB (used)

    Budget
    Budget
    Graphics memory
    ≈24 GB
    RAM
    ≈64 GB
    Components
    • GPU: RTX 3090 24 GB
    • CPU: any 6–8-core desktop CPU
    • Motherboard: any ATX with a PCIe x16 slot
    • RAM: 32–64 GB DDR4/DDR5
    • Power supply: 750–850 W
    • PCIe: 1× x16
    Runs
    Notes
    The cheapest way to 24 GB of fast VRAM. Small MoE models fly; dense ~27–32B models fit at Q4 with a moderate context. The Gemma 3 27B figure is our estimate: 936 GB/s × 55–75 % efficiency ÷ 16.5 GB of weights.
  • 12–24 GB GPU + 64–128 GB RAM (MoE offload)

    Budget
    Budget
    Graphics memory
    ≈24 GB
    RAM
    ≈64 GB
    Components
    • GPU: RTX 3090 24 GB / RTX 4070 12 GB / RTX 3080 Ti 12 GB
    • CPU: Core i5-12600K / Core Ultra 7 265K class
    • Motherboard: desktop board, both memory channels populated
    • RAM: 64–128 GB DDR5 (XMP/EXPO on) or DDR4-3600
    • Power supply: 750–850 W
    • PCIe: 1× x16
    Notes
    llama.cpp --n-cpu-moe N keeps attention and the shared layers on the GPU and moves the experts of N layers to system RAM. Works well only for MoE models; RAM speed matters — the 4070 build ran at 10–11 tok/s until XMP was turned on.
  • Two RTX 3090 (48 GB)

    Budget
    Mid-range
    Graphics memory
    ≈48 GB
    RAM
    ≈64 GB
    Components
    • GPU: 2× RTX 3090 24 GB
    • CPU: desktop CPU with x8/x8 bifurcation, or HEDT
    • Motherboard: two x16-size slots spaced for 3-slot cards (x8/x8)
    • RAM: 64 GB
    • Power supply: 1000–1200 W
    • PCIe: 2× x8 PCIe 4.0; NVLink optional
    Notes
    48 GB holds a dense 70B model at Q4_K_M (≈42.5 GB) with a short-to-medium context. llama.cpp splits the layers between the cards and runs them one after another, so speed is that of one 3090 reading the whole model; NVLink and x8 lanes barely matter here (they do for vLLM tensor parallel).
  • Used EPYC server, 8-channel DDR4 (CPU only)

    Budget
    Mid-range
    Graphics memory
    ≈0 GB
    RAM
    ≈512 GB
    Price
    $2,000
    Components
    • CPU: AMD EPYC 7002/7003 (Rome/Milan), e.g. EPYC 7702
    • Motherboard: ASRock Rack ROMED8-2T / Supermicro H12SSL-i / Gigabyte MZ32-AR0
    • RAM: 256–512 GB DDR4-3200 RDIMM, all 8 channels populated
    • Power supply: 750–1000 W
    • PCIe: 5–7× x16 PCIe 4.0 — room to add GPUs later
    Runs
    Notes
    Huge memory for little money: 8 channels of DDR4-3200 give 204.8 GB/s per socket. Only MoE models are practical — speed follows the active parameters. Fill every channel; a second socket adds NUMA trouble and little speed in llama.cpp (--numa distribute). The gpt-oss-120b figure is an estimate scaled by memory bandwidth from a DDR5 EPYC run. The source build cost about $2000 (prices of January 2025).
  • AMD Strix Halo mini-PC, 128 GB

    Budget
    Mid-range
    Graphics memory
    ≈96 GB
    RAM
    ≈128 GB
    Components
    • GPU: Radeon 8060S (integrated)
    • CPU: AMD Ryzen AI Max+ 395
    • Motherboard: Framework Desktop / other Strix Halo mini-PCs
    • RAM: 128 GB LPDDR5X-8000 unified (≈96 GB+ usable as VRAM)
    • Power supply: built-in
    Notes
    Quiet, ~120 W, the whole model in unified memory. Bandwidth (256 GB/s) is about a quarter of a 3090, so dense 70B models are slow; large MoE models with few active parameters are its sweet spot.
  • Three RTX 3090 (72 GB)

    Budget
    High-end
    Graphics memory
    ≈72 GB
    RAM
    ≈64 GB
    Components
    • GPU: 3× RTX 3090 24 GB
    • CPU: AMD EPYC 7002/7003 or Threadripper (enough PCIe lanes)
    • Motherboard: ASRock Rack ROMED8-2T / Supermicro H12SSL-i, risers for 3-slot cards
    • RAM: 64–128 GB
    • Power supply: 1200–1600 W (or two PSUs)
    • PCIe: 3× x8–x16 PCIe 4.0
    Runs
    Notes
    gpt-oss-120b (≈65 GB) fits fully in 72 GB of VRAM with a long context: 73 tok/s at 12.8k tokens of context, 41 tok/s at 93.7k. The measured machine was rented, so the platform here is our suggestion: a server board gives every card enough lanes. Three 350 W cards need a big PSU; many owners power-limit them.
  • NVIDIA DGX Spark, 128 GB

    Budget
    High-end
    Graphics memory
    ≈128 GB
    RAM
    ≈128 GB
    Price
    $3,999
    Components
    • GPU: NVIDIA GB10 (integrated Blackwell)
    • CPU: Arm CPU of the GB10
    • Motherboard: NVIDIA DGX Spark
    • RAM: 128 GB LPDDR5X unified
    • Power supply: built-in
    Runs
    Notes
    A CUDA box with 128 GB of unified memory at 273 GB/s: the same speed class as Strix Halo, with the NVIDIA software stack. Early llama.cpp builds gave ~35 tok/s on gpt-oss-120b; the figure is after the November 2025 update (llama-bench tg128).
  • Mac Studio, M2/M3 Ultra

    Budget
    High-end
    Graphics memory
    ≈192 GB
    RAM
    ≈192 GB
    Components
    • GPU: Apple M2 Ultra 76-core / M3 Ultra 80-core GPU
    • CPU: Apple M2 Ultra / M3 Ultra
    • Motherboard: Mac Studio
    • RAM: 192–512 GB unified (800–819 GB/s)
    • Power supply: built-in
    Notes
    The simplest way to run the biggest models at home: the M3 Ultra goes up to 512 GB of unified memory, enough for DeepSeek-class MoE models at 4 bits. Prompt processing is much slower than on NVIDIA GPUs. For very large models raise the GPU memory limit with sysctl iogpu.wired_limit_mb.

Measured speeds link to their source; ≈ marks an estimate from the memory bandwidth. Prices are approximate.

Updated · Sources: github.com , www.hardware-corner.net , carteakey.dev , github.crookster.org , hardware-corner.net , digitalspaceport.com , quozul.dev

Embed on your site

Paste this code where the infographic should appear: it updates by itself. Free to use — keep the attribution link.

JSON Our data: CC BY 4.0 with a link to the source; third-party data keeps the licence of its source.

Three budget groups

The cards group complete PCs for local models by budget. The budget group starts with one 16 GB graphics card, a used RTX 3090, and a GPU paired with plenty of RAM for MoE offload. Mid-range adds two RTX 3090, a used EPYC server with 8-channel DDR4 and an AMD Strix Halo mini-PC. High-end covers three RTX 3090, NVIDIA DGX Spark and a Mac Studio with M2/M3 Ultra. Each card lists graphics memory, RAM, an approximate price, the components down to the power supply and PCIe lanes, and the models it runs with their speed.

Measured or estimated

A speed without a mark was measured: it links to the source — usually a llama.cpp benchmark thread or a public report — and shows its date. A value marked ≈ is our estimate from the memory bandwidth. A measurement belongs to its runtime version and settings, so your machine may land somewhat higher or lower.

Choosing in four steps

  1. Model size first. Decide which models you want to run; that sets the memory. The hardware fit chart shows what each size needs.
  2. Then speed. Generation speed follows memory bandwidth. Graphics cards are fastest while the model fits in their memory; unified-memory machines hold bigger models but generate slower.
  3. Then the architecture of the model. CPU servers hold huge MoE models for little money, but only MoE is practical there, because each token reads just the active parameters.
  4. Finally noise, power and price. Several 3090s are loud and need a large power supply; a mini-PC or a Mac is quiet but slower per token.

Two things people get wrong

More cards do not mean more speed. With llama.cpp the layers are split between the GPUs and run one after another, so a second or third RTX 3090 lets you load a bigger model while the speed stays close to that of one card. The second point is MoE offload: llama.cpp's --n-cpu-moe keeps the shared layers on the GPU and moves the experts to system RAM. It works well only with MoE models, and the RAM speed matters.

Prices are approximate and change quickly, especially for used cards. To run a model on any of these machines you also need an app — see local AI apps — and a model file in the right quantization, explained in types of AI models.