Langsung ke konten
fedi.software

Rakitan PC untuk model AI lokal, 20B hingga 120B — 2026

Dari satu kartu grafis 16 GB hingga tiga RTX 3090, server EPYC, dan mesin bermemori terpadu: apa yang dijalankan tiap rakitan dan seberapa cepat, lengkap dengan sumber setiap pengukuran.

  • One 16 GB GPU

    Anggaran
    Hemat
    Memori grafis
    ≈16 GB
    RAM
    ≈32 GB
    Komponen
    • GPU: RTX 5060 Ti 16 GB / RTX 4060 Ti 16 GB
    • CPU: any 6–8-core desktop CPU
    • Motherboard: any ATX/mATX with a PCIe x16 slot
    • RAM: 32 GB DDR4/DDR5
    • Catu daya: 550–650 W
    • PCIe: 1× x16 (x8 electrical is enough)
    Menjalankan
    Catatan
    gpt-oss-20b (≈12 GB) fits entirely in VRAM with room for a long context. Qwen3-30B-A3B Q4_K_M (≈18.6 GB) does not fit: a few expert layers go to system RAM with --n-cpu-moe, still usable because only ~3B parameters are active. Measured on an RTX 5060 Ti 16 GB (llama-bench tg128).
  • One RTX 3090 24 GB (used)

    Anggaran
    Hemat
    Memori grafis
    ≈24 GB
    RAM
    ≈64 GB
    Komponen
    • GPU: RTX 3090 24 GB
    • CPU: any 6–8-core desktop CPU
    • Motherboard: any ATX with a PCIe x16 slot
    • RAM: 32–64 GB DDR4/DDR5
    • Catu daya: 750–850 W
    • PCIe: 1× x16
    Menjalankan
    Catatan
    The cheapest way to 24 GB of fast VRAM. Small MoE models fly; dense ~27–32B models fit at Q4 with a moderate context. The Gemma 3 27B figure is our estimate: 936 GB/s × 55–75 % efficiency ÷ 16.5 GB of weights.
  • 12–24 GB GPU + 64–128 GB RAM (MoE offload)

    Anggaran
    Hemat
    Memori grafis
    ≈24 GB
    RAM
    ≈64 GB
    Komponen
    • GPU: RTX 3090 24 GB / RTX 4070 12 GB / RTX 3080 Ti 12 GB
    • CPU: Core i5-12600K / Core Ultra 7 265K class
    • Motherboard: desktop board, both memory channels populated
    • RAM: 64–128 GB DDR5 (XMP/EXPO on) or DDR4-3600
    • Catu daya: 750–850 W
    • PCIe: 1× x16
    Catatan
    llama.cpp --n-cpu-moe N keeps attention and the shared layers on the GPU and moves the experts of N layers to system RAM. Works well only for MoE models; RAM speed matters — the 4070 build ran at 10–11 tok/s until XMP was turned on.
  • Two RTX 3090 (48 GB)

    Anggaran
    Menengah
    Memori grafis
    ≈48 GB
    RAM
    ≈64 GB
    Komponen
    • GPU: 2× RTX 3090 24 GB
    • CPU: desktop CPU with x8/x8 bifurcation, or HEDT
    • Motherboard: two x16-size slots spaced for 3-slot cards (x8/x8)
    • RAM: 64 GB
    • Catu daya: 1000–1200 W
    • PCIe: 2× x8 PCIe 4.0; NVLink optional
    Catatan
    48 GB holds a dense 70B model at Q4_K_M (≈42.5 GB) with a short-to-medium context. llama.cpp splits the layers between the cards and runs them one after another, so speed is that of one 3090 reading the whole model; NVLink and x8 lanes barely matter here (they do for vLLM tensor parallel).
  • Used EPYC server, 8-channel DDR4 (CPU only)

    Anggaran
    Menengah
    Memori grafis
    ≈0 GB
    RAM
    ≈512 GB
    Harga
    US$2.000
    Komponen
    • CPU: AMD EPYC 7002/7003 (Rome/Milan), e.g. EPYC 7702
    • Motherboard: ASRock Rack ROMED8-2T / Supermicro H12SSL-i / Gigabyte MZ32-AR0
    • RAM: 256–512 GB DDR4-3200 RDIMM, all 8 channels populated
    • Catu daya: 750–1000 W
    • PCIe: 5–7× x16 PCIe 4.0 — room to add GPUs later
    Menjalankan
    Catatan
    Huge memory for little money: 8 channels of DDR4-3200 give 204.8 GB/s per socket. Only MoE models are practical — speed follows the active parameters. Fill every channel; a second socket adds NUMA trouble and little speed in llama.cpp (--numa distribute). The gpt-oss-120b figure is an estimate scaled by memory bandwidth from a DDR5 EPYC run. The source build cost about $2000 (prices of January 2025).
  • AMD Strix Halo mini-PC, 128 GB

    Anggaran
    Menengah
    Memori grafis
    ≈96 GB
    RAM
    ≈128 GB
    Komponen
    • GPU: Radeon 8060S (integrated)
    • CPU: AMD Ryzen AI Max+ 395
    • Motherboard: Framework Desktop / other Strix Halo mini-PCs
    • RAM: 128 GB LPDDR5X-8000 unified (≈96 GB+ usable as VRAM)
    • Catu daya: built-in
    Menjalankan
    Catatan
    Quiet, ~120 W, the whole model in unified memory. Bandwidth (256 GB/s) is about a quarter of a 3090, so dense 70B models are slow; large MoE models with few active parameters are its sweet spot.
  • Three RTX 3090 (72 GB)

    Anggaran
    Kelas atas
    Memori grafis
    ≈72 GB
    RAM
    ≈64 GB
    Komponen
    • GPU: 3× RTX 3090 24 GB
    • CPU: AMD EPYC 7002/7003 or Threadripper (enough PCIe lanes)
    • Motherboard: ASRock Rack ROMED8-2T / Supermicro H12SSL-i, risers for 3-slot cards
    • RAM: 64–128 GB
    • Catu daya: 1200–1600 W (or two PSUs)
    • PCIe: 3× x8–x16 PCIe 4.0
    Menjalankan
    Catatan
    gpt-oss-120b (≈65 GB) fits fully in 72 GB of VRAM with a long context: 73 tok/s at 12.8k tokens of context, 41 tok/s at 93.7k. The measured machine was rented, so the platform here is our suggestion: a server board gives every card enough lanes. Three 350 W cards need a big PSU; many owners power-limit them.
  • NVIDIA DGX Spark, 128 GB

    Anggaran
    Kelas atas
    Memori grafis
    ≈128 GB
    RAM
    ≈128 GB
    Harga
    US$3.999
    Komponen
    • GPU: NVIDIA GB10 (integrated Blackwell)
    • CPU: Arm CPU of the GB10
    • Motherboard: NVIDIA DGX Spark
    • RAM: 128 GB LPDDR5X unified
    • Catu daya: built-in
    Menjalankan
    Catatan
    A CUDA box with 128 GB of unified memory at 273 GB/s: the same speed class as Strix Halo, with the NVIDIA software stack. Early llama.cpp builds gave ~35 tok/s on gpt-oss-120b; the figure is after the November 2025 update (llama-bench tg128).
  • Mac Studio, M2/M3 Ultra

    Anggaran
    Kelas atas
    Memori grafis
    ≈192 GB
    RAM
    ≈192 GB
    Komponen
    • GPU: Apple M2 Ultra 76-core / M3 Ultra 80-core GPU
    • CPU: Apple M2 Ultra / M3 Ultra
    • Motherboard: Mac Studio
    • RAM: 192–512 GB unified (800–819 GB/s)
    • Catu daya: built-in
    Catatan
    The simplest way to run the biggest models at home: the M3 Ultra goes up to 512 GB of unified memory, enough for DeepSeek-class MoE models at 4 bits. Prompt processing is much slower than on NVIDIA GPUs. For very large models raise the GPU memory limit with sysctl iogpu.wired_limit_mb.

Kecepatan terukur menautkan ke sumbernya; ≈ menandai perkiraan dari bandwidth memori. Harga bersifat perkiraan.

Diperbarui · Sumber: github.com , www.hardware-corner.net , carteakey.dev , github.crookster.org , hardware-corner.net , digitalspaceport.com , quozul.dev

Sematkan di situs Anda

Tempel kode ini di tempat infografis akan muncul: diperbarui dengan sendirinya. Gratis digunakan — pertahankan tautan atribusi.

JSON Data kami: CC BY 4.0 dengan tautan ke sumber; data pihak ketiga tetap memakai lisensi sumbernya.

Bagaimana rakitan dikelompokkan?

Kartu-kartu ini mengelompokkan PC lengkap untuk model lokal menurut anggaran. Kelompok «Hemat» dimulai dari satu kartu grafis 16 GB, satu RTX 3090 bekas, dan GPU yang dipasangkan dengan RAM besar untuk offload MoE. «Menengah» menambahkan dua RTX 3090, server EPYC bekas dengan DDR4 8 kanal, dan mini-PC AMD Strix Halo. «Kelas atas» mencakup tiga RTX 3090, NVIDIA DGX Spark, dan Mac Studio dengan M2/M3 Ultra. Setiap kartu mencantumkan memori grafis, RAM, perkiraan harga, komponen sampai catu daya dan jalur PCIe, serta model yang dijalankan beserta kecepatannya.

Kecepatannya hasil ukur atau perkiraan?

Kecepatan tanpa tanda adalah hasil pengukuran: tertaut ke sumbernya, biasanya utas benchmark llama.cpp atau laporan publik, dan menampilkan tanggalnya. Nilai bertanda ≈ adalah perkiraan kami dari bandwidth memori. Setiap pengukuran berlaku untuk versi runtime dan pengaturannya sendiri, jadi hasil di mesin Anda bisa sedikit lebih tinggi atau lebih rendah.

Mulai memilih dari mana?

  • Ukuran model. Tentukan model yang ingin dijalankan, karena itulah yang menetapkan kebutuhan memori. Grafik kecocokan perangkat keras menunjukkan kebutuhan tiap ukuran.
  • Kecepatan. Kecepatan generasi mengikuti bandwidth memori. Kartu grafis cepat selama model muat di memorinya; mesin memori terpadu (Mac, Strix Halo, DGX Spark) menampung model lebih besar tetapi menghasilkan teks lebih lambat.
  • Arsitektur model. Server CPU menampung model MoE raksasa dengan biaya rendah, tetapi di sana hanya MoE yang praktis, karena tiap token hanya membaca parameter aktif.
  • Kebisingan, daya, dan harga. Beberapa kartu 3090 itu bising dan butuh catu daya besar; mini-PC atau Mac senyap tetapi lebih lambat per token.

Apakah lebih banyak kartu berarti lebih cepat?

Tidak. Di llama.cpp lapisan model dibagi ke beberapa GPU dan dijalankan bergiliran, sehingga RTX 3090 kedua atau ketiga memungkinkan Anda memuat model yang lebih besar, sementara kecepatannya tetap mendekati satu kartu. Kesalahpahaman kedua menyangkut offload MoE: opsi --n-cpu-moe menyimpan lapisan bersama di GPU dan memindahkan para expert ke RAM sistem. Cara ini hanya bekerja baik untuk model MoE, dan kecepatan RAM ikut menentukan.

Seberapa akurat harganya?

Harga bersifat perkiraan dan cepat berubah, terutama untuk kartu bekas. Untuk menjalankan model, Anda juga memerlukan aplikasi — lihat aplikasi AI lokal — dan file model dengan kuantisasi yang tepat, seperti dijelaskan di jenis model AI.