Siirry sisältöön
fedi.software

PC-kokoonpanot paikallisille tekoälymalleille, 20B–120B — 2026

Yhdestä 16 GB:n näytönohjaimesta kolmeen RTX 3090:een, EPYC-palvelimeen ja yhtenäisen muistin koneisiin: mitä kukin kokoonpano ajaa ja kuinka nopeasti, jokaisen mittauksen lähteineen.

  • One 16 GB GPU

    Budjetti
    Edullinen
    Näytönohjaimen muisti
    ≈16 GB
    RAM
    ≈32 GB
    Komponentit
    • GPU: RTX 5060 Ti 16 GB / RTX 4060 Ti 16 GB
    • CPU: any 6–8-core desktop CPU
    • Emolevy: any ATX/mATX with a PCIe x16 slot
    • RAM: 32 GB DDR4/DDR5
    • Virtalähde: 550–650 W
    • PCIe: 1× x16 (x8 electrical is enough)
    Ajaa
    Huomautukset
    gpt-oss-20b (≈12 GB) fits entirely in VRAM with room for a long context. Qwen3-30B-A3B Q4_K_M (≈18.6 GB) does not fit: a few expert layers go to system RAM with --n-cpu-moe, still usable because only ~3B parameters are active. Measured on an RTX 5060 Ti 16 GB (llama-bench tg128).
  • One RTX 3090 24 GB (used)

    Budjetti
    Edullinen
    Näytönohjaimen muisti
    ≈24 GB
    RAM
    ≈64 GB
    Komponentit
    • GPU: RTX 3090 24 GB
    • CPU: any 6–8-core desktop CPU
    • Emolevy: any ATX with a PCIe x16 slot
    • RAM: 32–64 GB DDR4/DDR5
    • Virtalähde: 750–850 W
    • PCIe: 1× x16
    Ajaa
    Huomautukset
    The cheapest way to 24 GB of fast VRAM. Small MoE models fly; dense ~27–32B models fit at Q4 with a moderate context. The Gemma 3 27B figure is our estimate: 936 GB/s × 55–75 % efficiency ÷ 16.5 GB of weights.
  • 12–24 GB GPU + 64–128 GB RAM (MoE offload)

    Budjetti
    Edullinen
    Näytönohjaimen muisti
    ≈24 GB
    RAM
    ≈64 GB
    Komponentit
    • GPU: RTX 3090 24 GB / RTX 4070 12 GB / RTX 3080 Ti 12 GB
    • CPU: Core i5-12600K / Core Ultra 7 265K class
    • Emolevy: desktop board, both memory channels populated
    • RAM: 64–128 GB DDR5 (XMP/EXPO on) or DDR4-3600
    • Virtalähde: 750–850 W
    • PCIe: 1× x16
    Huomautukset
    llama.cpp --n-cpu-moe N keeps attention and the shared layers on the GPU and moves the experts of N layers to system RAM. Works well only for MoE models; RAM speed matters — the 4070 build ran at 10–11 tok/s until XMP was turned on.
  • Two RTX 3090 (48 GB)

    Budjetti
    Keskiluokka
    Näytönohjaimen muisti
    ≈48 GB
    RAM
    ≈64 GB
    Komponentit
    • GPU: 2× RTX 3090 24 GB
    • CPU: desktop CPU with x8/x8 bifurcation, or HEDT
    • Emolevy: two x16-size slots spaced for 3-slot cards (x8/x8)
    • RAM: 64 GB
    • Virtalähde: 1000–1200 W
    • PCIe: 2× x8 PCIe 4.0; NVLink optional
    Huomautukset
    48 GB holds a dense 70B model at Q4_K_M (≈42.5 GB) with a short-to-medium context. llama.cpp splits the layers between the cards and runs them one after another, so speed is that of one 3090 reading the whole model; NVLink and x8 lanes barely matter here (they do for vLLM tensor parallel).
  • Used EPYC server, 8-channel DDR4 (CPU only)

    Budjetti
    Keskiluokka
    Näytönohjaimen muisti
    ≈0 GB
    RAM
    ≈512 GB
    Hinta
    2 000 $
    Komponentit
    • CPU: AMD EPYC 7002/7003 (Rome/Milan), e.g. EPYC 7702
    • Emolevy: ASRock Rack ROMED8-2T / Supermicro H12SSL-i / Gigabyte MZ32-AR0
    • RAM: 256–512 GB DDR4-3200 RDIMM, all 8 channels populated
    • Virtalähde: 750–1000 W
    • PCIe: 5–7× x16 PCIe 4.0 — room to add GPUs later
    Ajaa
    Huomautukset
    Huge memory for little money: 8 channels of DDR4-3200 give 204.8 GB/s per socket. Only MoE models are practical — speed follows the active parameters. Fill every channel; a second socket adds NUMA trouble and little speed in llama.cpp (--numa distribute). The gpt-oss-120b figure is an estimate scaled by memory bandwidth from a DDR5 EPYC run. The source build cost about $2000 (prices of January 2025).
  • AMD Strix Halo mini-PC, 128 GB

    Budjetti
    Keskiluokka
    Näytönohjaimen muisti
    ≈96 GB
    RAM
    ≈128 GB
    Komponentit
    • GPU: Radeon 8060S (integrated)
    • CPU: AMD Ryzen AI Max+ 395
    • Emolevy: Framework Desktop / other Strix Halo mini-PCs
    • RAM: 128 GB LPDDR5X-8000 unified (≈96 GB+ usable as VRAM)
    • Virtalähde: built-in
    Ajaa
    Huomautukset
    Quiet, ~120 W, the whole model in unified memory. Bandwidth (256 GB/s) is about a quarter of a 3090, so dense 70B models are slow; large MoE models with few active parameters are its sweet spot.
  • Three RTX 3090 (72 GB)

    Budjetti
    Huippuluokka
    Näytönohjaimen muisti
    ≈72 GB
    RAM
    ≈64 GB
    Komponentit
    • GPU: 3× RTX 3090 24 GB
    • CPU: AMD EPYC 7002/7003 or Threadripper (enough PCIe lanes)
    • Emolevy: ASRock Rack ROMED8-2T / Supermicro H12SSL-i, risers for 3-slot cards
    • RAM: 64–128 GB
    • Virtalähde: 1200–1600 W (or two PSUs)
    • PCIe: 3× x8–x16 PCIe 4.0
    Ajaa
    Huomautukset
    gpt-oss-120b (≈65 GB) fits fully in 72 GB of VRAM with a long context: 73 tok/s at 12.8k tokens of context, 41 tok/s at 93.7k. The measured machine was rented, so the platform here is our suggestion: a server board gives every card enough lanes. Three 350 W cards need a big PSU; many owners power-limit them.
  • NVIDIA DGX Spark, 128 GB

    Budjetti
    Huippuluokka
    Näytönohjaimen muisti
    ≈128 GB
    RAM
    ≈128 GB
    Hinta
    3 999 $
    Komponentit
    • GPU: NVIDIA GB10 (integrated Blackwell)
    • CPU: Arm CPU of the GB10
    • Emolevy: NVIDIA DGX Spark
    • RAM: 128 GB LPDDR5X unified
    • Virtalähde: built-in
    Ajaa
    Huomautukset
    A CUDA box with 128 GB of unified memory at 273 GB/s: the same speed class as Strix Halo, with the NVIDIA software stack. Early llama.cpp builds gave ~35 tok/s on gpt-oss-120b; the figure is after the November 2025 update (llama-bench tg128).
  • Mac Studio, M2/M3 Ultra

    Budjetti
    Huippuluokka
    Näytönohjaimen muisti
    ≈192 GB
    RAM
    ≈192 GB
    Komponentit
    • GPU: Apple M2 Ultra 76-core / M3 Ultra 80-core GPU
    • CPU: Apple M2 Ultra / M3 Ultra
    • Emolevy: Mac Studio
    • RAM: 192–512 GB unified (800–819 GB/s)
    • Virtalähde: built-in
    Huomautukset
    The simplest way to run the biggest models at home: the M3 Ultra goes up to 512 GB of unified memory, enough for DeepSeek-class MoE models at 4 bits. Prompt processing is much slower than on NVIDIA GPUs. For very large models raise the GPU memory limit with sysctl iogpu.wired_limit_mb.

Mitatut nopeudet linkittävät lähteeseensä; ≈ merkitsee arviota muistin kaistanleveydestä. Hinnat ovat suuntaa-antavia.

Päivitetty · Lähteet: github.com , www.hardware-corner.net , carteakey.dev , github.crookster.org , hardware-corner.net , digitalspaceport.com , quozul.dev

Upota sivustollesi

Liitä tämä koodi kohtaan, johon infografiikan halutaan tulevan: se päivittyy itsestään. Vapaasti käytettävissä — säilytä lähdelinkki.

JSON Omat tietomme: CC BY 4.0 ja linkki lähteeseen; kolmansien osapuolten tiedoilla on lähteensä lisenssi.

Mitä kokoonpanoja kortit esittelevät?

Kortit ovat valmiita PC-koneita 20B–120B-kokoisille paikallisille malleille kolmessa hintaluokassa. ”Edullinen” sisältää yhden 16 GB:n näytönohjaimen, yhden käytetyn RTX 3090:n tai GPU:n ja runsaasti RAM-muistia MoE-siirtoa varten. ”Keskiluokassa” ovat kaksi RTX 3090:tä, käytetty EPYC-palvelin 8-kanavaisella DDR4-muistilla ja AMD Strix Halo -minitietokone. ”Huippuluokka” kattaa kolme RTX 3090:tä, NVIDIA DGX Sparkin ja Mac Studion M2/M3 Ultralla. Jokainen kortti kertoo näytönohjaimen muistin, RAM-muistin, likimääräisen hinnan, osat virtalähdettä ja PCIe-linjoja myöten sekä mallit, joita kone ajaa, ja niiden nopeuden.

Onko nopeudet mitattu?

Merkitsemätön nopeus on mitattu: sillä on päivämäärä ja linkki lähteeseen, yleensä llama.cpp:n vertailuketjuun tai julkiseen raporttiin. ≈-merkitty arvo on oma arviomme muistikaistan perusteella. Mittaus liittyy tiettyyn ajoympäristön versioon ja asetuksiin, joten oma koneesi voi päätyä hieman ylemmäs tai alemmas.

Miten valitsen?

  1. Mallien koko. Päätä ensin, mitä malleja haluat ajaa, sillä se määrää muistin. Mikä avoin malli sopii laitteistollesi näyttää kunkin koon tarpeen.
  2. Nopeus. Tuottonopeus seuraa muistikaistaa. Näytönohjaimet ovat nopeita, kunhan malli mahtuu niiden muistiin; yhteisen muistin koneet, kuten Mac, Strix Halo ja DGX Spark, kantavat isompia malleja mutta tuottavat hitaammin.
  3. Arkkitehtuuri. CPU-palvelin pitää valtavia malleja muistissa halvalla, mutta siellä vain MoE on käytännöllinen, koska jokainen token lukee vain aktiiviset parametrit.
  4. Melu, sähkö ja hinta. Useampi 3090 pitää ääntä ja vaatii järeän virtalähteen; minitietokone tai Mac on hiljainen mutta tokenia kohden hitaampi.

Nopeuttaako lisänäytönohjain?

Ei llama.cpp:llä. Kerrokset jaetaan GPU:iden kesken ja ajetaan peräkkäin, joten toinen tai kolmas RTX 3090 mahdollistaa isomman mallin, mutta nopeus pysyy lähellä yhden kortin tasoa.

Mikä on MoE-siirto?

Asetuksella --n-cpu-moe llama.cpp pitää jaetut kerrokset GPU:lla ja siirtää asiantuntijat järjestelmän RAM-muistiin. Tämä toimii hyvin vain MoE-malleilla, ja RAM-muistin nopeus ratkaisee paljon. Hinnat ovat suuntaa antavia ja muuttuvat nopeasti, etenkin käytettyjen korttien. Ohjelmat löydät sivulta paikalliset tekoälysovellukset, ja oikean mallitiedoston valintaa selittää tekoälymallien tyypit.