ข้ามไปยังเนื้อหา
fedi.software

ประเภทของโมเดล AI: แผนที่และอภิธานศัพท์

LLM หรือ VLM โอเพนเวตหรือโอเพนซอร์ส base, instruct, reasoning, distilled, MoE, GGUF และควอนไทเซชัน — คำในชื่อโมเดลหมายถึงอะไร พร้อมตัวอย่างสด

ประเภทโมเดล

ความเปิดกว้าง

สถาปัตยกรรม

รูปแบบไฟล์

ควอนไทเซชัน

กลุ่ม ความหมาย ตัวอย่าง
LLMlarge language model · text model ประเภทโมเดล A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal ประเภทโมเดล An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights ความเปิดกว้าง The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition ความเปิดกว้าง The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model ความเปิดกว้าง No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base รุ่นย่อย The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct รุ่นย่อย The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking รุ่นย่อย Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model รุ่นย่อย Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen รุ่นย่อย A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic ไม่มีตัวกรองความปลอดภัย รุ่นย่อย A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model สถาปัตยกรรม Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts สถาปัตยกรรม The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active สถาปัตยกรรม In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 รูปแบบไฟล์ The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp รูปแบบไฟล์ A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 รูปแบบไฟล์ Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community รูปแบบไฟล์ Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 ควอนไทเซชัน GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL ควอนไทเซชัน Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 ควอนไทเซชัน Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

โมเดลที่ทำเครื่องหมาย “ไม่มีตัวกรองความปลอดภัย” (uncensored / abliterated) แสดงไว้เพื่อการใช้งานส่วนตัวและการวิจัยเท่านั้น: ไม่มีข้อจำกัดด้านความปลอดภัย และผู้ใช้เป็นผู้รับผิดชอบการใช้งาน

อัปเดตเมื่อ · แหล่งที่มา: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

ฝังบนเว็บไซต์ของคุณ

วางโค้ดนี้ตรงที่ต้องการให้อินโฟกราฟิกปรากฏ: อัปเดตเองอัตโนมัติ ใช้ได้ฟรี — โปรดคงลิงก์อ้างอิงไว้

JSON ข้อมูลของเรา: CC BY 4.0 พร้อมลิงก์ไปยังแหล่งที่มา ส่วนข้อมูลของบุคคลที่สามใช้สัญญาอนุญาตของแหล่งที่มานั้น

อภิธานศัพท์หกกลุ่ม

คำที่พบในชื่อโมเดลแบ่งออกเป็นหกกลุ่ม และทุกคำมีตัวอย่างโมเดลจริงที่ลิงก์ไปยังหน้าเปรียบเทียบโมเดล

  • ประเภทโมเดล LLM ที่ทำงานกับข้อความ และ VLM ที่อ่านภาพได้ด้วย
  • ความเปิดกว้าง โอเพนเวต, Open Source AI และแบบกรรมสิทธิ์
  • รุ่นย่อย base, instruct, reasoning, coder, distill และ abliterated / uncensored
  • สถาปัตยกรรม แบบหนาแน่น (dense), MoE และพารามิเตอร์ที่ทำงานอยู่
  • รูปแบบไฟล์ Safetensors, GGUF, AWQ / GPTQ / EXL และ MLX
  • ควอนไทเซชัน K-quants อย่าง Q4_K_M, IQ และ Unsloth Dynamic, FP8 / MXFP4

อ่านชื่อโมเดลทีละส่วน

อ่านจากหน้าไปหลัง เริ่มจากตระกูลและขนาด ถ้าเป็น MoE จะตามด้วยพารามิเตอร์ที่ทำงานอยู่ เช่น -A3B ต่อด้วยรุ่นย่อยอย่าง -Instruct หรือ -it และท้ายสุดคือรูปแบบไฟล์กับควอนไทซ์ของไฟล์ที่ดาวน์โหลด โมเดลกลั่น (distill) คือโมเดลเล็กที่ฝึกจากคำตอบของโมเดลใหญ่ ไม่ใช่ตัวโมเดลใหญ่เอง

ขั้นตอนเลือกโมเดลสำหรับแชตในเครื่อง

  1. เลือกรุ่นย่อย Instruct (หรือ reasoning)
  2. ดูว่าเครื่องของคุณมีหน่วยความจำเท่าไร
  3. เลือกควอนไทซ์ที่ใส่ได้ Q4_K_M เป็นจุดสมดุลที่นิยมระหว่างขนาดกับคุณภาพ Q8_0 แทบไม่เสียคุณภาพ ส่วนต่ำกว่า Q3 คุณภาพลดลงอย่างเห็นได้ชัด
  4. กำหนดความยาวคอนเท็กซ์ที่ต้องการ
  5. ตรวจความเร็วที่รับได้ในโมเดลเปิดตัวไหนใส่ฮาร์ดแวร์ของคุณได้

โอเพนเวตไม่เท่ากับโอเพนซอร์ส

โอเพนเวต (open weights) อาจมาพร้อมสัญญาอนุญาตที่จำกัดการใช้งาน และข้อมูลที่ใช้ฝึกมักไม่เปิดเผย เกณฑ์ที่เราอ้างอิงคือ OSI Open Source AI Definition 1.0 แหล่งข้อมูลอื่นคือการ์ดโมเดลและไฟล์ตั้งค่าบน Hugging Face และเอกสารของ llama.cpp

โมเดล abliterated / uncensored คือโมเดลที่ถูกตัดพฤติกรรมการปฏิเสธคำขอออก เราแสดงไว้อย่างเป็นกลางในฐานะคำศัพท์เท่านั้น ไม่ได้แนะนำให้ใช้ ใช้ได้เพื่อการส่วนตัวและการวิจัยเท่านั้น และผู้ใช้เป็นผู้รับผิดชอบการใช้งานเอง