로컬 AI 모델용 PC 구성, 20B~120B — 2026
16 GB 그래픽 카드 한 장부터 RTX 3090 세 장, EPYC 서버, 통합 메모리 머신까지: 각 구성에서 무엇이 얼마나 빠르게 실행되는지, 모든 측정의 출처와 함께 보여 줍니다.
-
One 16 GB GPU
- 예산
- 보급형
- 그래픽 메모리
- ≈16 GB
- RAM
- ≈32 GB
- 부품
- GPU: RTX 5060 Ti 16 GB / RTX 4060 Ti 16 GB
- CPU: any 6–8-core desktop CPU
- 메인보드: any ATX/mATX with a PCIe x16 slot
- RAM: 32 GB DDR4/DDR5
- 파워 서플라이: 550–650 W
- PCIe: 1× x16 (x8 electrical is enough)
- 실행 가능
- gpt-oss-20b MXFP4 (RTX 5060 Ti 16 GB) MXFP4 112 tok/s [실측]
- Qwen3-30B-A3B Q4_K_M, experts partly in RAM Q4_K_M [≈ 메모리 대역폭 기반 추정]
- 비고
- gpt-oss-20b (≈12 GB) fits entirely in VRAM with room for a long context. Qwen3-30B-A3B Q4_K_M (≈18.6 GB) does not fit: a few expert layers go to system RAM with --n-cpu-moe, still usable because only ~3B parameters are active. Measured on an RTX 5060 Ti 16 GB (llama-bench tg128).
-
One RTX 3090 24 GB (used)
- 예산
- 보급형
- 그래픽 메모리
- ≈24 GB
- RAM
- ≈64 GB
- 부품
- GPU: RTX 3090 24 GB
- CPU: any 6–8-core desktop CPU
- 메인보드: any ATX with a PCIe x16 slot
- RAM: 32–64 GB DDR4/DDR5
- 파워 서플라이: 750–850 W
- PCIe: 1× x16
- 실행 가능
- gpt-oss-20b MXFP4 MXFP4 162 tok/s [실측]
- Qwen3-30B-A3B Q4_K Q4_K_M 154 tok/s @4k [실측]
- Gemma 3 27B Q4_K_M Q4_K_M ≈31–42 tok/s [≈ 메모리 대역폭 기반 추정]
- 비고
- The cheapest way to 24 GB of fast VRAM. Small MoE models fly; dense ~27–32B models fit at Q4 with a moderate context. The Gemma 3 27B figure is our estimate: 936 GB/s × 55–75 % efficiency ÷ 16.5 GB of weights.
-
12–24 GB GPU + 64–128 GB RAM (MoE offload)
- 예산
- 보급형
- 그래픽 메모리
- ≈24 GB
- RAM
- ≈64 GB
- 부품
- GPU: RTX 3090 24 GB / RTX 4070 12 GB / RTX 3080 Ti 12 GB
- CPU: Core i5-12600K / Core Ultra 7 265K class
- 메인보드: desktop board, both memory channels populated
- RAM: 64–128 GB DDR5 (XMP/EXPO on) or DDR4-3600
- 파워 서플라이: 750–850 W
- PCIe: 1× x16
- 실행 가능
- gpt-oss-120b MXFP4 — RTX 3090 + 64 GB DDR5-5200, --n-cpu-moe 24–26 MXFP4 26–28 tok/s [실측]
- gpt-oss-120b MXFP4 — RTX 4070 12 GB + 64 GB DDR5 MXFP4 25–28 tok/s [실측]
- gpt-oss-120b MXFP4 — RTX 3080 Ti 12 GB + 128 GB DDR4-3600, --n-cpu-moe 36 MXFP4 18.0–22 tok/s [실측]
- 비고
- llama.cpp --n-cpu-moe N keeps attention and the shared layers on the GPU and moves the experts of N layers to system RAM. Works well only for MoE models; RAM speed matters — the 4070 build ran at 10–11 tok/s until XMP was turned on.
-
Two RTX 3090 (48 GB)
- 예산
- 중급형
- 그래픽 메모리
- ≈48 GB
- RAM
- ≈64 GB
- 부품
- GPU: 2× RTX 3090 24 GB
- CPU: desktop CPU with x8/x8 bifurcation, or HEDT
- 메인보드: two x16-size slots spaced for 3-slot cards (x8/x8)
- RAM: 64 GB
- 파워 서플라이: 1000–1200 W
- PCIe: 2× x8 PCIe 4.0; NVLink optional
- 실행 가능
- Llama 3 70B Q4_K_M (same size as Llama 3.3 70B) Q4_K_M 16.3 tok/s [실측]
- 비고
- 48 GB holds a dense 70B model at Q4_K_M (≈42.5 GB) with a short-to-medium context. llama.cpp splits the layers between the cards and runs them one after another, so speed is that of one 3090 reading the whole model; NVLink and x8 lanes barely matter here (they do for vLLM tensor parallel).
-
Used EPYC server, 8-channel DDR4 (CPU only)
- 예산
- 중급형
- 그래픽 메모리
- ≈0 GB
- RAM
- ≈512 GB
- 가격
- US$2,000
- 부품
- CPU: AMD EPYC 7002/7003 (Rome/Milan), e.g. EPYC 7702
- 메인보드: ASRock Rack ROMED8-2T / Supermicro H12SSL-i / Gigabyte MZ32-AR0
- RAM: 256–512 GB DDR4-3200 RDIMM, all 8 channels populated
- 파워 서플라이: 750–1000 W
- PCIe: 5–7× x16 PCIe 4.0 — room to add GPUs later
- 실행 가능
- DeepSeek R1 671B Q4 — EPYC 7702 + 512 GB DDR4-2400, MZ32-AR0 Q4 3.5–4.2 tok/s [실측]
- gpt-oss-120b MXFP4 MXFP4 ≈12.0–16.0 tok/s [≈ 메모리 대역폭 기반 추정]
- 비고
- Huge memory for little money: 8 channels of DDR4-3200 give 204.8 GB/s per socket. Only MoE models are practical — speed follows the active parameters. Fill every channel; a second socket adds NUMA trouble and little speed in llama.cpp (--numa distribute). The gpt-oss-120b figure is an estimate scaled by memory bandwidth from a DDR5 EPYC run. The source build cost about $2000 (prices of January 2025).
-
AMD Strix Halo mini-PC, 128 GB
- 예산
- 중급형
- 그래픽 메모리
- ≈96 GB
- RAM
- ≈128 GB
- 부품
- GPU: Radeon 8060S (integrated)
- CPU: AMD Ryzen AI Max+ 395
- 메인보드: Framework Desktop / other Strix Halo mini-PCs
- RAM: 128 GB LPDDR5X-8000 unified (≈96 GB+ usable as VRAM)
- 파워 서플라이: built-in
- 실행 가능
- gpt-oss-120b MXFP4 (ROCm) MXFP4 50 tok/s [실측]
- 비고
- Quiet, ~120 W, the whole model in unified memory. Bandwidth (256 GB/s) is about a quarter of a 3090, so dense 70B models are slow; large MoE models with few active parameters are its sweet spot.
-
Three RTX 3090 (72 GB)
- 예산
- 고급형
- 그래픽 메모리
- ≈72 GB
- RAM
- ≈64 GB
- 부품
- GPU: 3× RTX 3090 24 GB
- CPU: AMD EPYC 7002/7003 or Threadripper (enough PCIe lanes)
- 메인보드: ASRock Rack ROMED8-2T / Supermicro H12SSL-i, risers for 3-slot cards
- RAM: 64–128 GB
- 파워 서플라이: 1200–1600 W (or two PSUs)
- PCIe: 3× x8–x16 PCIe 4.0
- 실행 가능
- gpt-oss-120b MXFP4, all in VRAM MXFP4 41–73 tok/s @12.8k→93.7k [실측]
- 비고
- gpt-oss-120b (≈65 GB) fits fully in 72 GB of VRAM with a long context: 73 tok/s at 12.8k tokens of context, 41 tok/s at 93.7k. The measured machine was rented, so the platform here is our suggestion: a server board gives every card enough lanes. Three 350 W cards need a big PSU; many owners power-limit them.
-
NVIDIA DGX Spark, 128 GB
- 예산
- 고급형
- 그래픽 메모리
- ≈128 GB
- RAM
- ≈128 GB
- 가격
- US$3,999
- 부품
- GPU: NVIDIA GB10 (integrated Blackwell)
- CPU: Arm CPU of the GB10
- 메인보드: NVIDIA DGX Spark
- RAM: 128 GB LPDDR5X unified
- 파워 서플라이: built-in
- 실행 가능
- gpt-oss-120b MXFP4 MXFP4 61 tok/s [실측]
- 비고
- A CUDA box with 128 GB of unified memory at 273 GB/s: the same speed class as Strix Halo, with the NVIDIA software stack. Early llama.cpp builds gave ~35 tok/s on gpt-oss-120b; the figure is after the November 2025 update (llama-bench tg128).
-
Mac Studio, M2/M3 Ultra
- 예산
- 고급형
- 그래픽 메모리
- ≈192 GB
- RAM
- ≈192 GB
- 부품
- GPU: Apple M2 Ultra 76-core / M3 Ultra 80-core GPU
- CPU: Apple M2 Ultra / M3 Ultra
- 메인보드: Mac Studio
- RAM: 192–512 GB unified (800–819 GB/s)
- 파워 서플라이: built-in
- 실행 가능
- gpt-oss-120b MXFP4 — M2 Ultra 192 GB MXFP4 80 tok/s [실측]
- gpt-oss-20b MXFP4 — M3 Ultra 512 GB MXFP4 116 tok/s [실측]
- 비고
- The simplest way to run the biggest models at home: the M3 Ultra goes up to 512 GB of unified memory, enough for DeepSeek-class MoE models at 4 bits. Prompt processing is much slower than on NVIDIA GPUs. For very large models raise the GPU memory limit with sysctl iogpu.wired_limit_mb.
실측 속도는 출처로 연결됩니다. ≈는 메모리 대역폭 기반 추정치입니다. 가격은 대략적인 값입니다.
업데이트 · 출처: github.com , www.hardware-corner.net , carteakey.dev , github.crookster.org , hardware-corner.net , digitalspaceport.com , quozul.dev
내 사이트에 삽입
인포그래픽을 표시할 위치에 이 코드를 붙여 넣으세요. 저절로 업데이트됩니다. 무료로 사용할 수 있으며, 출처 링크는 유지해 주세요.
세 가지 예산대
로컬 모델용 완성 PC 구성을 예산별로 묶은 카드입니다. ‘보급형’은 16 GB 그래픽 카드 한 장, 중고 RTX 3090 한 장, MoE 오프로드용으로 RAM을 넉넉히 단 GPU 구성입니다. ‘중급형’은 RTX 3090 두 장, 8채널 DDR4 중고 EPYC 서버, AMD Strix Halo 미니 PC입니다. ‘고급형’은 RTX 3090 세 장, NVIDIA DGX Spark, M2/M3 Ultra를 단 Mac Studio입니다. 카드마다 그래픽 메모리, RAM, 대략적인 가격, GPU부터 전원 공급 장치와 PCIe 레인까지의 부품, 실행되는 모델과 속도, 비고가 있습니다.
구성을 고르는 순서
- 돌리고 싶은 모델 크기를 먼저 정합니다. 그것이 필요한 메모리를 결정합니다.
- 받아들일 수 있는 속도, 곧 메모리 대역폭을 봅니다. 그래픽 카드는 모델이 그 메모리에 들어가는 동안 빠르고, Mac, Strix Halo, DGX Spark 같은 통합 메모리 머신은 더 큰 모델을 담지만 생성이 느립니다.
- CPU 서버는 거대한 MoE 모델을 싸게 담을 수 있으나 실용적인 것은 MoE뿐입니다.
- 마지막으로 소음, 전력, 가격을 따집니다.
실측과 추정
표시가 없는 속도는 실측값으로, 출처(주로 llama.cpp 벤치마크 스레드와 공개 보고서) 링크와 날짜가 붙어 있습니다. ≈ 표시는 메모리 대역폭으로 계산한 추정치입니다. 실측은 당시 런타임 버전과 설정의 결과라서 내 환경과 차이가 날 수 있습니다.
자주 하는 오해
카드를 늘린다고 빨라지지 않습니다. llama.cpp에서는 여러 카드가 모델을 나눠 담지만 차례로 처리하므로, 더 큰 모델을 올릴 수 있어도 속도는 카드 한 장일 때와 비슷합니다. MoE 오프로드(--n-cpu-moe)는 공유 레이어를 GPU에 두고 전문가(expert)를 시스템 RAM에 둡니다. MoE 모델에서만 잘 작동하며 RAM 속도가 결과를 좌우합니다. 가격은 대략적인 값이고 중고 카드는 상태에 따라 다릅니다. 실행용 앱은 로컬 AI 앱, 양자화 읽는 법은 AI 모델의 종류에서 볼 수 있습니다.