ข้ามไปยังเนื้อหา
fedi.software

สเปกพีซีสำหรับโมเดล AI ในเครื่อง 20B ถึง 120B — 2026

ตั้งแต่การ์ดจอ 16 GB ใบเดียวไปจนถึง RTX 3090 สามใบ เซิร์ฟเวอร์ EPYC และเครื่องที่ใช้หน่วยความจำแบบรวม: แต่ละสเปกรันอะไรได้และเร็วแค่ไหน พร้อมแหล่งที่มาของทุกการวัด

  • One 16 GB GPU

    งบประมาณ
    ประหยัด
    หน่วยความจำการ์ดจอ
    ≈16 GB
    RAM
    ≈32 GB
    ชิ้นส่วน
    • GPU: RTX 5060 Ti 16 GB / RTX 4060 Ti 16 GB
    • CPU: any 6–8-core desktop CPU
    • เมนบอร์ด: any ATX/mATX with a PCIe x16 slot
    • RAM: 32 GB DDR4/DDR5
    • พาวเวอร์ซัพพลาย: 550–650 W
    • PCIe: 1× x16 (x8 electrical is enough)
    รันได้
    หมายเหตุ
    gpt-oss-20b (≈12 GB) fits entirely in VRAM with room for a long context. Qwen3-30B-A3B Q4_K_M (≈18.6 GB) does not fit: a few expert layers go to system RAM with --n-cpu-moe, still usable because only ~3B parameters are active. Measured on an RTX 5060 Ti 16 GB (llama-bench tg128).
  • One RTX 3090 24 GB (used)

    งบประมาณ
    ประหยัด
    หน่วยความจำการ์ดจอ
    ≈24 GB
    RAM
    ≈64 GB
    ชิ้นส่วน
    • GPU: RTX 3090 24 GB
    • CPU: any 6–8-core desktop CPU
    • เมนบอร์ด: any ATX with a PCIe x16 slot
    • RAM: 32–64 GB DDR4/DDR5
    • พาวเวอร์ซัพพลาย: 750–850 W
    • PCIe: 1× x16
    รันได้
    หมายเหตุ
    The cheapest way to 24 GB of fast VRAM. Small MoE models fly; dense ~27–32B models fit at Q4 with a moderate context. The Gemma 3 27B figure is our estimate: 936 GB/s × 55–75 % efficiency ÷ 16.5 GB of weights.
  • 12–24 GB GPU + 64–128 GB RAM (MoE offload)

    งบประมาณ
    ประหยัด
    หน่วยความจำการ์ดจอ
    ≈24 GB
    RAM
    ≈64 GB
    ชิ้นส่วน
    • GPU: RTX 3090 24 GB / RTX 4070 12 GB / RTX 3080 Ti 12 GB
    • CPU: Core i5-12600K / Core Ultra 7 265K class
    • เมนบอร์ด: desktop board, both memory channels populated
    • RAM: 64–128 GB DDR5 (XMP/EXPO on) or DDR4-3600
    • พาวเวอร์ซัพพลาย: 750–850 W
    • PCIe: 1× x16
    หมายเหตุ
    llama.cpp --n-cpu-moe N keeps attention and the shared layers on the GPU and moves the experts of N layers to system RAM. Works well only for MoE models; RAM speed matters — the 4070 build ran at 10–11 tok/s until XMP was turned on.
  • Two RTX 3090 (48 GB)

    งบประมาณ
    ระดับกลาง
    หน่วยความจำการ์ดจอ
    ≈48 GB
    RAM
    ≈64 GB
    ชิ้นส่วน
    • GPU: 2× RTX 3090 24 GB
    • CPU: desktop CPU with x8/x8 bifurcation, or HEDT
    • เมนบอร์ด: two x16-size slots spaced for 3-slot cards (x8/x8)
    • RAM: 64 GB
    • พาวเวอร์ซัพพลาย: 1000–1200 W
    • PCIe: 2× x8 PCIe 4.0; NVLink optional
    หมายเหตุ
    48 GB holds a dense 70B model at Q4_K_M (≈42.5 GB) with a short-to-medium context. llama.cpp splits the layers between the cards and runs them one after another, so speed is that of one 3090 reading the whole model; NVLink and x8 lanes barely matter here (they do for vLLM tensor parallel).
  • Used EPYC server, 8-channel DDR4 (CPU only)

    งบประมาณ
    ระดับกลาง
    หน่วยความจำการ์ดจอ
    ≈0 GB
    RAM
    ≈512 GB
    ราคา
    US$2,000
    ชิ้นส่วน
    • CPU: AMD EPYC 7002/7003 (Rome/Milan), e.g. EPYC 7702
    • เมนบอร์ด: ASRock Rack ROMED8-2T / Supermicro H12SSL-i / Gigabyte MZ32-AR0
    • RAM: 256–512 GB DDR4-3200 RDIMM, all 8 channels populated
    • พาวเวอร์ซัพพลาย: 750–1000 W
    • PCIe: 5–7× x16 PCIe 4.0 — room to add GPUs later
    รันได้
    หมายเหตุ
    Huge memory for little money: 8 channels of DDR4-3200 give 204.8 GB/s per socket. Only MoE models are practical — speed follows the active parameters. Fill every channel; a second socket adds NUMA trouble and little speed in llama.cpp (--numa distribute). The gpt-oss-120b figure is an estimate scaled by memory bandwidth from a DDR5 EPYC run. The source build cost about $2000 (prices of January 2025).
  • AMD Strix Halo mini-PC, 128 GB

    งบประมาณ
    ระดับกลาง
    หน่วยความจำการ์ดจอ
    ≈96 GB
    RAM
    ≈128 GB
    ชิ้นส่วน
    • GPU: Radeon 8060S (integrated)
    • CPU: AMD Ryzen AI Max+ 395
    • เมนบอร์ด: Framework Desktop / other Strix Halo mini-PCs
    • RAM: 128 GB LPDDR5X-8000 unified (≈96 GB+ usable as VRAM)
    • พาวเวอร์ซัพพลาย: built-in
    รันได้
    หมายเหตุ
    Quiet, ~120 W, the whole model in unified memory. Bandwidth (256 GB/s) is about a quarter of a 3090, so dense 70B models are slow; large MoE models with few active parameters are its sweet spot.
  • Three RTX 3090 (72 GB)

    งบประมาณ
    ระดับสูง
    หน่วยความจำการ์ดจอ
    ≈72 GB
    RAM
    ≈64 GB
    ชิ้นส่วน
    • GPU: 3× RTX 3090 24 GB
    • CPU: AMD EPYC 7002/7003 or Threadripper (enough PCIe lanes)
    • เมนบอร์ด: ASRock Rack ROMED8-2T / Supermicro H12SSL-i, risers for 3-slot cards
    • RAM: 64–128 GB
    • พาวเวอร์ซัพพลาย: 1200–1600 W (or two PSUs)
    • PCIe: 3× x8–x16 PCIe 4.0
    รันได้
    หมายเหตุ
    gpt-oss-120b (≈65 GB) fits fully in 72 GB of VRAM with a long context: 73 tok/s at 12.8k tokens of context, 41 tok/s at 93.7k. The measured machine was rented, so the platform here is our suggestion: a server board gives every card enough lanes. Three 350 W cards need a big PSU; many owners power-limit them.
  • NVIDIA DGX Spark, 128 GB

    งบประมาณ
    ระดับสูง
    หน่วยความจำการ์ดจอ
    ≈128 GB
    RAM
    ≈128 GB
    ราคา
    US$3,999
    ชิ้นส่วน
    • GPU: NVIDIA GB10 (integrated Blackwell)
    • CPU: Arm CPU of the GB10
    • เมนบอร์ด: NVIDIA DGX Spark
    • RAM: 128 GB LPDDR5X unified
    • พาวเวอร์ซัพพลาย: built-in
    รันได้
    หมายเหตุ
    A CUDA box with 128 GB of unified memory at 273 GB/s: the same speed class as Strix Halo, with the NVIDIA software stack. Early llama.cpp builds gave ~35 tok/s on gpt-oss-120b; the figure is after the November 2025 update (llama-bench tg128).
  • Mac Studio, M2/M3 Ultra

    งบประมาณ
    ระดับสูง
    หน่วยความจำการ์ดจอ
    ≈192 GB
    RAM
    ≈192 GB
    ชิ้นส่วน
    • GPU: Apple M2 Ultra 76-core / M3 Ultra 80-core GPU
    • CPU: Apple M2 Ultra / M3 Ultra
    • เมนบอร์ด: Mac Studio
    • RAM: 192–512 GB unified (800–819 GB/s)
    • พาวเวอร์ซัพพลาย: built-in
    หมายเหตุ
    The simplest way to run the biggest models at home: the M3 Ultra goes up to 512 GB of unified memory, enough for DeepSeek-class MoE models at 4 bits. Prompt processing is much slower than on NVIDIA GPUs. For very large models raise the GPU memory limit with sysctl iogpu.wired_limit_mb.

ความเร็วที่วัดได้มีลิงก์ไปยังแหล่งที่มา ส่วน ≈ คือค่าประมาณจากแบนด์วิดท์หน่วยความจำ ราคาเป็นราคาโดยประมาณ

อัปเดตเมื่อ · แหล่งที่มา: github.com , www.hardware-corner.net , carteakey.dev , github.crookster.org , hardware-corner.net , digitalspaceport.com , quozul.dev

ฝังบนเว็บไซต์ของคุณ

วางโค้ดนี้ตรงที่ต้องการให้อินโฟกราฟิกปรากฏ: อัปเดตเองอัตโนมัติ ใช้ได้ฟรี — โปรดคงลิงก์อ้างอิงไว้

JSON ข้อมูลของเรา: CC BY 4.0 พร้อมลิงก์ไปยังแหล่งที่มา ส่วนข้อมูลของบุคคลที่สามใช้สัญญาอนุญาตของแหล่งที่มานั้น

สามระดับงบประมาณ

การ์ดในหน้านี้รวบรวมสเปกพีซีสำเร็จสำหรับโมเดลในเครื่อง แยกตามงบ กลุ่ม “ประหยัด” มีการ์ดจอ 16 GB หนึ่งใบ RTX 3090 มือสองหนึ่งใบ และ GPU ที่จับคู่กับ RAM จำนวนมากสำหรับย้าย MoE ไปที่ RAM กลุ่ม “ระดับกลาง” มี RTX 3090 สองใบ เซิร์ฟเวอร์ EPYC มือสองที่ใช้ DDR4 แบบ 8 แชนเนล และมินิพีซี AMD Strix Halo กลุ่ม “ระดับสูง” มี RTX 3090 สามใบ NVIDIA DGX Spark และ Mac Studio ที่ใช้ M2/M3 Ultra การ์ดแต่ละใบบอกหน่วยความจำการ์ดจอ RAM ราคาโดยประมาณ ชิ้นส่วนตั้งแต่ GPU จนถึงพาวเวอร์ซัพพลายและเลน PCIe โมเดลที่รันได้พร้อมความเร็ว และหมายเหตุ

วัดจริงหรือประมาณ

ความเร็วที่ไม่มีเครื่องหมายคือค่าที่วัดจริง มีลิงก์ไปยังแหล่งที่มา ซึ่งส่วนใหญ่เป็นกระทู้ทดสอบของ llama.cpp หรือรายงานสาธารณะ พร้อมวันที่ ค่าที่มีเครื่องหมาย ≈ คือค่าประมาณจากแบนด์วิดท์หน่วยความจำ ผลวัดจริงขึ้นกับเวอร์ชันรันไทม์และการตั้งค่าในตอนนั้น ผลบนเครื่องของคุณจึงอาจสูงหรือต่ำกว่าเล็กน้อย

เลือกสเปกอย่างไร

เริ่มจากขนาดโมเดลที่อยากรัน เพราะขนาดเป็นตัวกำหนดหน่วยความจำ จากนั้นดูความเร็วที่รับได้ ซึ่งก็คือแบนด์วิดท์หน่วยความจำ แล้วค่อยพิจารณาเสียง ไฟ และราคา การ์ดจอเร็วตราบใดที่โมเดลยังอยู่ในหน่วยความจำของการ์ด เครื่องที่ใช้หน่วยความจำรวมอย่าง Mac, Strix Halo และ DGX Spark รับโมเดลใหญ่กว่าได้แต่สร้างข้อความช้ากว่า เซิร์ฟเวอร์ CPU รับโมเดล MoE ขนาดมหึมาได้ในราคาถูก แต่ใช้งานได้จริงเฉพาะ MoE

สองเรื่องที่มักเข้าใจผิด

  • เพิ่มการ์ดไม่ได้ทำให้เร็วขึ้น ใน llama.cpp การ์ดหลายใบแบ่งกันเก็บโมเดลแต่ทำงานต่อกันทีละใบ จึงโหลดโมเดลใหญ่ขึ้นได้ ขณะที่ความเร็วยังใกล้กับการ์ดใบเดียว
  • การย้าย MoE ไปที่ RAM (--n-cpu-moe) เก็บเลเยอร์ที่ใช้ร่วมกันไว้บน GPU และย้ายเอ็กซ์เพิร์ตไปไว้ใน RAM ของระบบ วิธีนี้ได้ผลดีกับโมเดล MoE เท่านั้น และความเร็วของ RAM มีผลมาก

ราคาเป็นค่าโดยประมาณและเปลี่ยนได้ โดยเฉพาะการ์ดมือสองที่สภาพต่างกัน ไม่ว่าจะเลือกเครื่องไหน คุณยังต้องมีแอปสำหรับรันโมเดล (ดูเครื่องมือ AI ในเครื่อง) และไฟล์ที่ควอนไทซ์อย่างเหมาะสม วิธีอ่านควอนไทซ์อธิบายไว้ในประเภทของโมเดล AI