跳到主要內容
fedi.software

AI 模型類型:地圖與名詞表

LLM 或 VLM、開放權重或開源、base、instruct、reasoning、蒸餾、MoE、GGUF 與量化 — 模型名稱中這些字詞的意思,附即時範例。

分組 意思 範例
LLMlarge language model · text model 模型類型 A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal 模型類型 An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights 開放程度 The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition 開放程度 The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model 開放程度 No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base 版本變體 The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct 版本變體 The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking 版本變體 Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model 版本變體 Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen 版本變體 A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic 無安全過濾 版本變體 A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model 架構 Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts 架構 The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active 架構 In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 檔案格式 The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp 檔案格式 A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 檔案格式 Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community 檔案格式 Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 量化 GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL 量化 Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 量化 Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

標示為「無安全過濾」的模型(uncensored / abliterated)僅供個人與研究用途:它們沒有安全限制,使用責任由使用者自行承擔。

更新日期 · 來源: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

嵌入你的網站

將此程式碼貼在資訊圖表要顯示的位置:它會自動更新。可免費使用 — 請保留出處連結。

JSON 我們的資料:CC BY 4.0,須附來源連結;第三方資料沿用其來源的授權。

名詞表分成哪幾組?

模型名稱裡常見的詞分成六組,每個術語都附上實際模型範例,並連到模型比較:

  • 模型類型:處理文字的 LLM、也能看圖的 VLM;
  • 開放程度:開放權重、Open Source AI、專有;
  • 版本變體:base、instruct、reasoning、coder、distill、abliterated / uncensored;
  • 架構:稠密(dense)、MoE、活躍參數;
  • 檔案格式:Safetensors、GGUF、AWQ / GPTQ / EXL、MLX;
  • 量化:Q4_K_M 等 K-quants、IQ 與 Unsloth Dynamic、FP8 / MXFP4。

模型名稱要怎麼讀?

從前往後拆:先是系列與規模,MoE 會接著標活躍參數(例如 -A3B),然後是 -Instruct 或 -it 之類的變體,最後是下載檔案的格式與量化。蒸餾(distill)模型是用大模型的回答訓練出來的小模型,並不是大模型本身。

本機聊天該挑哪種模型?

先選 Instruct(或 reasoning)變體,再確認手上的記憶體,挑放得下的量化,接著決定需要的上下文長度,最後到哪個開放模型適合您的硬體確認能接受的速度。Q4_K_M 是大小與品質之間常見的平衡點,Q8_0 幾乎無損,低於 Q3 品質會明顯下降。

開放權重就是開源嗎?

不是。開放權重(open weights)可能附帶限制用途的授權,訓練資料通常也不公開;判斷標準可參考 OSI 的 Open Source AI Definition 1.0。來源還包括 Hugging Face 的模型卡與設定檔,以及 llama.cpp 文件。

abliterated / uncensored 是什麼?

這類模型移除了拒答行為。名詞表只以中立方式列出,不作推薦;僅限個人與研究用途,使用責任由使用者自行承擔。