Vai al contenuto
fedi.software

Configurazioni PC per modelli di IA locali, da 20B a 120B — 2026

Da una singola scheda video da 16 GB a tre RTX 3090, un server EPYC e macchine a memoria unificata: cosa esegue ogni configurazione e a che velocità, con la fonte di ogni misurazione.

  • One 16 GB GPU

    Budget
    Economica
    Memoria video
    ≈16 GB
    RAM
    ≈32 GB
    Componenti
    • GPU: RTX 5060 Ti 16 GB / RTX 4060 Ti 16 GB
    • CPU: any 6–8-core desktop CPU
    • Scheda madre: any ATX/mATX with a PCIe x16 slot
    • RAM: 32 GB DDR4/DDR5
    • Alimentatore: 550–650 W
    • PCIe: 1× x16 (x8 electrical is enough)
    Esegue
    Note
    gpt-oss-20b (≈12 GB) fits entirely in VRAM with room for a long context. Qwen3-30B-A3B Q4_K_M (≈18.6 GB) does not fit: a few expert layers go to system RAM with --n-cpu-moe, still usable because only ~3B parameters are active. Measured on an RTX 5060 Ti 16 GB (llama-bench tg128).
  • One RTX 3090 24 GB (used)

    Budget
    Economica
    Memoria video
    ≈24 GB
    RAM
    ≈64 GB
    Componenti
    • GPU: RTX 3090 24 GB
    • CPU: any 6–8-core desktop CPU
    • Scheda madre: any ATX with a PCIe x16 slot
    • RAM: 32–64 GB DDR4/DDR5
    • Alimentatore: 750–850 W
    • PCIe: 1× x16
    Esegue
    Note
    The cheapest way to 24 GB of fast VRAM. Small MoE models fly; dense ~27–32B models fit at Q4 with a moderate context. The Gemma 3 27B figure is our estimate: 936 GB/s × 55–75 % efficiency ÷ 16.5 GB of weights.
  • 12–24 GB GPU + 64–128 GB RAM (MoE offload)

    Budget
    Economica
    Memoria video
    ≈24 GB
    RAM
    ≈64 GB
    Componenti
    • GPU: RTX 3090 24 GB / RTX 4070 12 GB / RTX 3080 Ti 12 GB
    • CPU: Core i5-12600K / Core Ultra 7 265K class
    • Scheda madre: desktop board, both memory channels populated
    • RAM: 64–128 GB DDR5 (XMP/EXPO on) or DDR4-3600
    • Alimentatore: 750–850 W
    • PCIe: 1× x16
    Note
    llama.cpp --n-cpu-moe N keeps attention and the shared layers on the GPU and moves the experts of N layers to system RAM. Works well only for MoE models; RAM speed matters — the 4070 build ran at 10–11 tok/s until XMP was turned on.
  • Two RTX 3090 (48 GB)

    Budget
    Fascia media
    Memoria video
    ≈48 GB
    RAM
    ≈64 GB
    Componenti
    • GPU: 2× RTX 3090 24 GB
    • CPU: desktop CPU with x8/x8 bifurcation, or HEDT
    • Scheda madre: two x16-size slots spaced for 3-slot cards (x8/x8)
    • RAM: 64 GB
    • Alimentatore: 1000–1200 W
    • PCIe: 2× x8 PCIe 4.0; NVLink optional
    Note
    48 GB holds a dense 70B model at Q4_K_M (≈42.5 GB) with a short-to-medium context. llama.cpp splits the layers between the cards and runs them one after another, so speed is that of one 3090 reading the whole model; NVLink and x8 lanes barely matter here (they do for vLLM tensor parallel).
  • Used EPYC server, 8-channel DDR4 (CPU only)

    Budget
    Fascia media
    Memoria video
    ≈0 GB
    RAM
    ≈512 GB
    Prezzo
    2.000 USD
    Componenti
    • CPU: AMD EPYC 7002/7003 (Rome/Milan), e.g. EPYC 7702
    • Scheda madre: ASRock Rack ROMED8-2T / Supermicro H12SSL-i / Gigabyte MZ32-AR0
    • RAM: 256–512 GB DDR4-3200 RDIMM, all 8 channels populated
    • Alimentatore: 750–1000 W
    • PCIe: 5–7× x16 PCIe 4.0 — room to add GPUs later
    Esegue
    Note
    Huge memory for little money: 8 channels of DDR4-3200 give 204.8 GB/s per socket. Only MoE models are practical — speed follows the active parameters. Fill every channel; a second socket adds NUMA trouble and little speed in llama.cpp (--numa distribute). The gpt-oss-120b figure is an estimate scaled by memory bandwidth from a DDR5 EPYC run. The source build cost about $2000 (prices of January 2025).
  • AMD Strix Halo mini-PC, 128 GB

    Budget
    Fascia media
    Memoria video
    ≈96 GB
    RAM
    ≈128 GB
    Componenti
    • GPU: Radeon 8060S (integrated)
    • CPU: AMD Ryzen AI Max+ 395
    • Scheda madre: Framework Desktop / other Strix Halo mini-PCs
    • RAM: 128 GB LPDDR5X-8000 unified (≈96 GB+ usable as VRAM)
    • Alimentatore: built-in
    Esegue
    Note
    Quiet, ~120 W, the whole model in unified memory. Bandwidth (256 GB/s) is about a quarter of a 3090, so dense 70B models are slow; large MoE models with few active parameters are its sweet spot.
  • Three RTX 3090 (72 GB)

    Budget
    Fascia alta
    Memoria video
    ≈72 GB
    RAM
    ≈64 GB
    Componenti
    • GPU: 3× RTX 3090 24 GB
    • CPU: AMD EPYC 7002/7003 or Threadripper (enough PCIe lanes)
    • Scheda madre: ASRock Rack ROMED8-2T / Supermicro H12SSL-i, risers for 3-slot cards
    • RAM: 64–128 GB
    • Alimentatore: 1200–1600 W (or two PSUs)
    • PCIe: 3× x8–x16 PCIe 4.0
    Esegue
    Note
    gpt-oss-120b (≈65 GB) fits fully in 72 GB of VRAM with a long context: 73 tok/s at 12.8k tokens of context, 41 tok/s at 93.7k. The measured machine was rented, so the platform here is our suggestion: a server board gives every card enough lanes. Three 350 W cards need a big PSU; many owners power-limit them.
  • NVIDIA DGX Spark, 128 GB

    Budget
    Fascia alta
    Memoria video
    ≈128 GB
    RAM
    ≈128 GB
    Prezzo
    3.999 USD
    Componenti
    • GPU: NVIDIA GB10 (integrated Blackwell)
    • CPU: Arm CPU of the GB10
    • Scheda madre: NVIDIA DGX Spark
    • RAM: 128 GB LPDDR5X unified
    • Alimentatore: built-in
    Esegue
    Note
    A CUDA box with 128 GB of unified memory at 273 GB/s: the same speed class as Strix Halo, with the NVIDIA software stack. Early llama.cpp builds gave ~35 tok/s on gpt-oss-120b; the figure is after the November 2025 update (llama-bench tg128).
  • Mac Studio, M2/M3 Ultra

    Budget
    Fascia alta
    Memoria video
    ≈192 GB
    RAM
    ≈192 GB
    Componenti
    • GPU: Apple M2 Ultra 76-core / M3 Ultra 80-core GPU
    • CPU: Apple M2 Ultra / M3 Ultra
    • Scheda madre: Mac Studio
    • RAM: 192–512 GB unified (800–819 GB/s)
    • Alimentatore: built-in
    Note
    The simplest way to run the biggest models at home: the M3 Ultra goes up to 512 GB of unified memory, enough for DeepSeek-class MoE models at 4 bits. Prompt processing is much slower than on NVIDIA GPUs. For very large models raise the GPU memory limit with sysctl iogpu.wired_limit_mb.

Le velocità misurate rimandano alla fonte; ≈ indica una stima dalla larghezza di banda della memoria. I prezzi sono indicativi.

Aggiornato · Fonti: github.com , www.hardware-corner.net , carteakey.dev , github.crookster.org , hardware-corner.net , digitalspaceport.com , quozul.dev

Incorpora nel tuo sito

Incolla questo codice dove deve apparire l’infografica: si aggiorna da sola. Uso gratuito — mantieni il link di attribuzione.

JSON I nostri dati: CC BY 4.0 con link alla fonte; i dati di terzi mantengono la licenza della loro fonte.

Tre fasce di configurazioni

Queste schede descrivono PC completi per eseguire in casa modelli da 20B a 120B, raccolti in tre fasce di budget. «Economica» comprende una scheda video da 16 GB, una RTX 3090 usata oppure una GPU con molta RAM per l'offload MoE. «Fascia media» riunisce due RTX 3090, un server EPYC usato con DDR4 a 8 canali e un mini-PC AMD Strix Halo. «Fascia alta» va da tre RTX 3090 a NVIDIA DGX Spark fino al Mac Studio con M2/M3 Ultra.

Ogni scheda riporta memoria video, RAM, prezzo indicativo e componenti: GPU, CPU, scheda madre, RAM, alimentatore e linee PCIe. Seguono i modelli che la macchina esegue e a che velocità. Ogni velocità è misurata, con link alla fonte e data, oppure stimata dalla larghezza di banda della memoria e segnata con ≈.

Scegliere in tre passi

  1. Parti dalla dimensione dei modelli che vuoi usare: è lei a decidere la memoria.
  2. Poi la velocità che ti basta, legata alla larghezza di banda. Le schede video sono rapide finché il modello ci sta per intero; le macchine a memoria unificata come Mac, Strix Halo o DGX Spark contengono modelli più grandi ma generano più piano; i server CPU ospitano enormi modelli MoE a poco prezzo, ma lì è pratico solo il MoE.
  3. Solo alla fine rumore, consumi e prezzo.

Due errori comuni

Più schede non vuol dire più velocità. Con llama.cpp gli strati si dividono tra le GPU, che lavorano una dopo l'altra: una seconda o terza RTX 3090 ti permette di caricare un modello più grande, ma la velocità resta vicina a quella di una scheda sola. L'offload MoE (--n-cpu-moe) invece tiene gli strati condivisi sulla GPU e sposta gli esperti nella RAM di sistema; funziona bene, ma solo con i modelli MoE, e la velocità della RAM pesa parecchio.

I prezzi sono indicativi e cambiano, soprattutto sull'usato. Una misura è stata fatta con la propria versione del runtime e le proprie impostazioni, quindi il tuo risultato può essere diverso. Ti servono poi un programma tra le app di IA locale e un file del modello nella quantizzazione giusta, spiegata in tipi di modelli di IA.