AI 모델의 종류: 지도와 용어집
LLM과 VLM, 오픈 웨이트와 오픈 소스, base, instruct, reasoning, 증류, MoE, GGUF, 양자화 — 모델 이름 속 용어의 의미를 실제 예시와 함께 설명합니다.
모델 유형
공개 수준
아키텍처
파일 형식
양자화
| 그룹 | 의미 | 예시 | |
|---|---|---|---|
| LLMlarge language model · text model | 모델 유형 | A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it. | |
| VLMVL · vision-language · multimodal | 모델 유형 | An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights. | |
| Open weightsopen model · downloadable weights | 공개 수준 | The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed. | |
| Open Source AIOSAID · OSI definition | 공개 수준 | The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. | — |
| Proprietary (API only)closed · API model | 공개 수준 | No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. | — |
| Basepretrained · -Base | 변형 | The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. | — |
| Instruct-it · -Chat · -Instruct | 변형 | The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat. | |
| ReasoningThinking · R1 · -Thinking | 변형 | Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply. | |
| Coder-Coder · code model | 변형 | Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model. | |
| Distill-Distill · R1-Distill-Qwen | 변형 | A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. | — |
| Abliterated / uncensoreduncensored · abliterated · heretic 안전 필터 없음 | 변형 | A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. | — |
| Densedense model | 아키텍처 | Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token. | |
| MoEmixture of experts | 아키텍처 | The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM. | |
| Active parameters (-A3B)-A3B · -A22B · active | 아키텍처 | In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number. | |
| Safetensors (BF16)safetensors · BF16 · FP16 | 파일 형식 | The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. | — |
| GGUF.gguf · llama.cpp | 파일 형식 | A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). | — |
| AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 | 파일 형식 | Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. | — |
| MLXmlx-community | 파일 형식 | Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. | — |
| Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 | 양자화 | GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. | — |
| IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL | 양자화 | Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable. | |
| FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 | 양자화 | Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB. |
“안전 필터 없음”으로 표시된 모델(uncensored / abliterated)은 개인 및 연구 목적으로만 게재됩니다. 안전 제한이 없으며, 사용에 대한 책임은 사용자에게 있습니다.
업데이트 · 출처: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)
내 사이트에 삽입
인포그래픽을 표시할 위치에 이 코드를 붙여 넣으세요. 저절로 업데이트됩니다. 무료로 사용할 수 있으며, 출처 링크는 유지해 주세요.
용어집의 여섯 그룹
모델 이름에 나오는 말을 여섯 그룹으로 나눠 설명하고, 용어마다 실제 모델 예시를 모델 비교로 연결했습니다.
- 모델 유형: 텍스트를 다루는 LLM, 이미지도 읽는 VLM.
- 공개 수준: 오픈 웨이트, Open Source AI, 독점.
- 변형: base, instruct, reasoning, coder, distill, abliterated / uncensored.
- 아키텍처: 밀집(dense), MoE, 활성 파라미터.
- 파일 형식: Safetensors, GGUF, AWQ / GPTQ / EXL, MLX.
- 양자화: Q4_K_M 같은 K-quants, IQ와 Unsloth Dynamic, FP8 / MXFP4.
이름은 조각별로 읽습니다
앞에서부터 계열과 크기, MoE라면 -A3B 같은 활성 파라미터, 이어서 -Instruct나 -it 같은 변형, 끝으로 내려받을 파일의 형식과 양자화가 옵니다. 로컬 채팅용이라면 Instruct(또는 reasoning) 변형을 고르고, 보유 메모리를 확인한 뒤 들어가는 양자화, 필요한 컨텍스트, 받아들일 수 있는 속도 순으로 정합니다. 마지막 단계는 내 하드웨어에 맞는 오픈 모델에서 확인할 수 있습니다. Q4_K_M은 크기와 품질의 균형이 맞는 흔한 선택이고, Q8_0은 손실이 거의 없으며, Q3 아래로 내려가면 품질이 눈에 띄게 떨어집니다.
헷갈리기 쉬운 개념
오픈 웨이트(open weights)는 오픈 소스와 같은 말이 아닙니다. 가중치가 공개되어도 사용을 제한하는 라이선스가 붙을 수 있고, 학습 데이터는 대개 비공개입니다. 기준으로는 OSI Open Source AI Definition 1.0을 참고합니다. 증류(distill) 모델은 큰 모델의 답변으로 학습한 작은 모델이며, 큰 모델 자체가 아닙니다.
abliterated / uncensored 모델은 거부 동작을 제거한 모델입니다. 용어로서 중립적으로 실었을 뿐 권장하지 않으며, 개인적·연구 목적으로만 쓰이고 사용 책임은 사용자에게 있습니다. 출처는 Hugging Face 모델 카드와 설정 파일, llama.cpp 문서입니다.