본문으로 건너뛰기
fedi.software

로컬 AI 모델용 PC 구성, 20B~120B — 2026

16 GB 그래픽 카드 한 장부터 RTX 3090 세 장, EPYC 서버, 통합 메모리 머신까지: 각 구성에서 무엇이 얼마나 빠르게 실행되는지, 모든 측정의 출처와 함께 보여 줍니다.

  • One 16 GB GPU

    예산
    보급형
    그래픽 메모리
    ≈16 GB
    RAM
    ≈32 GB
    부품
    • GPU: RTX 5060 Ti 16 GB / RTX 4060 Ti 16 GB
    • CPU: any 6–8-core desktop CPU
    • 메인보드: any ATX/mATX with a PCIe x16 slot
    • RAM: 32 GB DDR4/DDR5
    • 파워 서플라이: 550–650 W
    • PCIe: 1× x16 (x8 electrical is enough)
    실행 가능
    비고
    gpt-oss-20b (≈12 GB) fits entirely in VRAM with room for a long context. Qwen3-30B-A3B Q4_K_M (≈18.6 GB) does not fit: a few expert layers go to system RAM with --n-cpu-moe, still usable because only ~3B parameters are active. Measured on an RTX 5060 Ti 16 GB (llama-bench tg128).
  • One RTX 3090 24 GB (used)

    예산
    보급형
    그래픽 메모리
    ≈24 GB
    RAM
    ≈64 GB
    부품
    • GPU: RTX 3090 24 GB
    • CPU: any 6–8-core desktop CPU
    • 메인보드: any ATX with a PCIe x16 slot
    • RAM: 32–64 GB DDR4/DDR5
    • 파워 서플라이: 750–850 W
    • PCIe: 1× x16
    실행 가능
    비고
    The cheapest way to 24 GB of fast VRAM. Small MoE models fly; dense ~27–32B models fit at Q4 with a moderate context. The Gemma 3 27B figure is our estimate: 936 GB/s × 55–75 % efficiency ÷ 16.5 GB of weights.
  • 12–24 GB GPU + 64–128 GB RAM (MoE offload)

    예산
    보급형
    그래픽 메모리
    ≈24 GB
    RAM
    ≈64 GB
    부품
    • GPU: RTX 3090 24 GB / RTX 4070 12 GB / RTX 3080 Ti 12 GB
    • CPU: Core i5-12600K / Core Ultra 7 265K class
    • 메인보드: desktop board, both memory channels populated
    • RAM: 64–128 GB DDR5 (XMP/EXPO on) or DDR4-3600
    • 파워 서플라이: 750–850 W
    • PCIe: 1× x16
    비고
    llama.cpp --n-cpu-moe N keeps attention and the shared layers on the GPU and moves the experts of N layers to system RAM. Works well only for MoE models; RAM speed matters — the 4070 build ran at 10–11 tok/s until XMP was turned on.
  • Two RTX 3090 (48 GB)

    예산
    중급형
    그래픽 메모리
    ≈48 GB
    RAM
    ≈64 GB
    부품
    • GPU: 2× RTX 3090 24 GB
    • CPU: desktop CPU with x8/x8 bifurcation, or HEDT
    • 메인보드: two x16-size slots spaced for 3-slot cards (x8/x8)
    • RAM: 64 GB
    • 파워 서플라이: 1000–1200 W
    • PCIe: 2× x8 PCIe 4.0; NVLink optional
    비고
    48 GB holds a dense 70B model at Q4_K_M (≈42.5 GB) with a short-to-medium context. llama.cpp splits the layers between the cards and runs them one after another, so speed is that of one 3090 reading the whole model; NVLink and x8 lanes barely matter here (they do for vLLM tensor parallel).
  • Used EPYC server, 8-channel DDR4 (CPU only)

    예산
    중급형
    그래픽 메모리
    ≈0 GB
    RAM
    ≈512 GB
    가격
    US$2,000
    부품
    • CPU: AMD EPYC 7002/7003 (Rome/Milan), e.g. EPYC 7702
    • 메인보드: ASRock Rack ROMED8-2T / Supermicro H12SSL-i / Gigabyte MZ32-AR0
    • RAM: 256–512 GB DDR4-3200 RDIMM, all 8 channels populated
    • 파워 서플라이: 750–1000 W
    • PCIe: 5–7× x16 PCIe 4.0 — room to add GPUs later
    실행 가능
    비고
    Huge memory for little money: 8 channels of DDR4-3200 give 204.8 GB/s per socket. Only MoE models are practical — speed follows the active parameters. Fill every channel; a second socket adds NUMA trouble and little speed in llama.cpp (--numa distribute). The gpt-oss-120b figure is an estimate scaled by memory bandwidth from a DDR5 EPYC run. The source build cost about $2000 (prices of January 2025).
  • AMD Strix Halo mini-PC, 128 GB

    예산
    중급형
    그래픽 메모리
    ≈96 GB
    RAM
    ≈128 GB
    부품
    • GPU: Radeon 8060S (integrated)
    • CPU: AMD Ryzen AI Max+ 395
    • 메인보드: Framework Desktop / other Strix Halo mini-PCs
    • RAM: 128 GB LPDDR5X-8000 unified (≈96 GB+ usable as VRAM)
    • 파워 서플라이: built-in
    실행 가능
    비고
    Quiet, ~120 W, the whole model in unified memory. Bandwidth (256 GB/s) is about a quarter of a 3090, so dense 70B models are slow; large MoE models with few active parameters are its sweet spot.
  • Three RTX 3090 (72 GB)

    예산
    고급형
    그래픽 메모리
    ≈72 GB
    RAM
    ≈64 GB
    부품
    • GPU: 3× RTX 3090 24 GB
    • CPU: AMD EPYC 7002/7003 or Threadripper (enough PCIe lanes)
    • 메인보드: ASRock Rack ROMED8-2T / Supermicro H12SSL-i, risers for 3-slot cards
    • RAM: 64–128 GB
    • 파워 서플라이: 1200–1600 W (or two PSUs)
    • PCIe: 3× x8–x16 PCIe 4.0
    실행 가능
    비고
    gpt-oss-120b (≈65 GB) fits fully in 72 GB of VRAM with a long context: 73 tok/s at 12.8k tokens of context, 41 tok/s at 93.7k. The measured machine was rented, so the platform here is our suggestion: a server board gives every card enough lanes. Three 350 W cards need a big PSU; many owners power-limit them.
  • NVIDIA DGX Spark, 128 GB

    예산
    고급형
    그래픽 메모리
    ≈128 GB
    RAM
    ≈128 GB
    가격
    US$3,999
    부품
    • GPU: NVIDIA GB10 (integrated Blackwell)
    • CPU: Arm CPU of the GB10
    • 메인보드: NVIDIA DGX Spark
    • RAM: 128 GB LPDDR5X unified
    • 파워 서플라이: built-in
    실행 가능
    비고
    A CUDA box with 128 GB of unified memory at 273 GB/s: the same speed class as Strix Halo, with the NVIDIA software stack. Early llama.cpp builds gave ~35 tok/s on gpt-oss-120b; the figure is after the November 2025 update (llama-bench tg128).
  • Mac Studio, M2/M3 Ultra

    예산
    고급형
    그래픽 메모리
    ≈192 GB
    RAM
    ≈192 GB
    부품
    • GPU: Apple M2 Ultra 76-core / M3 Ultra 80-core GPU
    • CPU: Apple M2 Ultra / M3 Ultra
    • 메인보드: Mac Studio
    • RAM: 192–512 GB unified (800–819 GB/s)
    • 파워 서플라이: built-in
    비고
    The simplest way to run the biggest models at home: the M3 Ultra goes up to 512 GB of unified memory, enough for DeepSeek-class MoE models at 4 bits. Prompt processing is much slower than on NVIDIA GPUs. For very large models raise the GPU memory limit with sysctl iogpu.wired_limit_mb.

실측 속도는 출처로 연결됩니다. ≈는 메모리 대역폭 기반 추정치입니다. 가격은 대략적인 값입니다.

업데이트 · 출처: github.com , www.hardware-corner.net , carteakey.dev , github.crookster.org , hardware-corner.net , digitalspaceport.com , quozul.dev

내 사이트에 삽입

인포그래픽을 표시할 위치에 이 코드를 붙여 넣으세요. 저절로 업데이트됩니다. 무료로 사용할 수 있으며, 출처 링크는 유지해 주세요.

JSON 자체 데이터: CC BY 4.0, 출처 링크 필요. 제3자 데이터는 출처의 라이선스를 따릅니다.

세 가지 예산대

로컬 모델용 완성 PC 구성을 예산별로 묶은 카드입니다. ‘보급형’은 16 GB 그래픽 카드 한 장, 중고 RTX 3090 한 장, MoE 오프로드용으로 RAM을 넉넉히 단 GPU 구성입니다. ‘중급형’은 RTX 3090 두 장, 8채널 DDR4 중고 EPYC 서버, AMD Strix Halo 미니 PC입니다. ‘고급형’은 RTX 3090 세 장, NVIDIA DGX Spark, M2/M3 Ultra를 단 Mac Studio입니다. 카드마다 그래픽 메모리, RAM, 대략적인 가격, GPU부터 전원 공급 장치와 PCIe 레인까지의 부품, 실행되는 모델과 속도, 비고가 있습니다.

구성을 고르는 순서

  1. 돌리고 싶은 모델 크기를 먼저 정합니다. 그것이 필요한 메모리를 결정합니다.
  2. 받아들일 수 있는 속도, 곧 메모리 대역폭을 봅니다. 그래픽 카드는 모델이 그 메모리에 들어가는 동안 빠르고, Mac, Strix Halo, DGX Spark 같은 통합 메모리 머신은 더 큰 모델을 담지만 생성이 느립니다.
  3. CPU 서버는 거대한 MoE 모델을 싸게 담을 수 있으나 실용적인 것은 MoE뿐입니다.
  4. 마지막으로 소음, 전력, 가격을 따집니다.

실측과 추정

표시가 없는 속도는 실측값으로, 출처(주로 llama.cpp 벤치마크 스레드와 공개 보고서) 링크와 날짜가 붙어 있습니다. ≈ 표시는 메모리 대역폭으로 계산한 추정치입니다. 실측은 당시 런타임 버전과 설정의 결과라서 내 환경과 차이가 날 수 있습니다.

자주 하는 오해

카드를 늘린다고 빨라지지 않습니다. llama.cpp에서는 여러 카드가 모델을 나눠 담지만 차례로 처리하므로, 더 큰 모델을 올릴 수 있어도 속도는 카드 한 장일 때와 비슷합니다. MoE 오프로드(--n-cpu-moe)는 공유 레이어를 GPU에 두고 전문가(expert)를 시스템 RAM에 둡니다. MoE 모델에서만 잘 작동하며 RAM 속도가 결과를 좌우합니다. 가격은 대략적인 값이고 중고 카드는 상태에 따라 다릅니다. 실행용 앱은 로컬 AI 앱, 양자화 읽는 법은 AI 모델의 종류에서 볼 수 있습니다.