Langsung ke konten
fedi.software

Jenis model AI: peta dan glosarium

LLM atau VLM, open weights atau sumber terbuka, base, instruct, reasoning, distilasi, MoE, GGUF, dan kuantisasi — arti kata-kata dalam nama model, dengan contoh langsung.

Grup Artinya Contoh
LLMlarge language model · text model Jenis model A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal Jenis model An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights Keterbukaan The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition Keterbukaan The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model Keterbukaan No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base Varian The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct Varian The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking Varian Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model Varian Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen Varian A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic Tanpa filter keamanan Varian A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model Arsitektur Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts Arsitektur The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active Arsitektur In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 Format file The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp Format file A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 Format file Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community Format file Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 Kuantisasi GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL Kuantisasi Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 Kuantisasi Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

Model yang bertanda “tanpa filter keamanan” (uncensored / abliterated) dicantumkan hanya untuk keperluan pribadi dan penelitian: model ini tidak memiliki pembatasan keamanan, dan tanggung jawab penggunaannya ada pada pengguna.

Diperbarui · Sumber: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

Sematkan di situs Anda

Tempel kode ini di tempat infografis akan muncul: diperbarui dengan sendirinya. Gratis digunakan — pertahankan tautan atribusi.

JSON Data kami: CC BY 4.0 dengan tautan ke sumber; data pihak ketiga tetap memakai lisensi sumbernya.

Membedah nama model

Nama seperti «Qwen3-30B-A3B-Instruct-GGUF Q4_K_M» memadatkan beberapa keputusan ke dalam satu baris. Pertama keluarga dan ukuran total; pada model MoE lalu parameter aktif (-A3B); kemudian varian (-Instruct atau -it); terakhir format file dan kuantisasi unduhan. Glosarium di halaman ini membagi kata-kata tersebut ke dalam enam kelompok, dengan contoh langsung di samping tiap istilah — klik sebuah contoh untuk membukanya di perbandingan model.

Enam kelompok secara singkat

  • Jenis model: LLM bekerja dengan teks, VLM juga menerima gambar.
  • Keterbukaan: open weights, definisi Open Source AI dari OSI yang lebih ketat, atau model proprietary yang hanya tersedia lewat aplikasi atau API.
  • Varian: base, Instruct, reasoning, coder, model hasil distilasi, dan versi komunitas yang penolakannya dihapus.
  • Arsitektur: dense lawan MoE, serta akhiran parameter aktif.
  • Format file: Safetensors, GGUF, format khusus GPU seperti AWQ, GPTQ, dan EXL, serta MLX untuk Mac.
  • Kuantisasi: K-quant seperti Q4_K_M, file IQ dan Unsloth Dynamic, FP8, dan MXFP4.

Dari istilah ke pilihan

Untuk asisten chat lokal, urutan keputusannya cukup tetap. Ambil varian Instruct atau reasoning, jangan model base. Lihat berapa memori yang Anda punya, lalu pilih kuantisasi yang muat: Q4_K_M adalah keseimbangan umum antara ukuran dan kualitas, Q8_0 hampir tanpa kehilangan, sedangkan di bawah Q3 jawaban menurun dengan jelas. Sisakan ruang untuk konteks yang dibutuhkan, baru setelah itu nilai apakah kecepatannya bisa diterima. Hitungan ini untuk model-model terkenal sudah dikerjakan oleh grafik kecocokan perangkat keras.

Dua salah kaprah

Open weights tidak sama dengan open source. File-nya bisa diunduh, tetapi lisensinya dapat membatasi penggunaan komersial atau meminta Anda menyetujui syarat tertentu, dan data pelatihannya biasanya tetap tertutup. Model hasil distilasi juga sering disalahpahami: itu model lebih kecil yang dilatih dari jawaban model yang lebih besar, bukan model besar itu sendiri, dan kemampuannya jelas lebih lemah.

Tentang model abliterated dan uncensored

Ini adalah modifikasi komunitas atas model terbuka yang perilaku menolaknya dihilangkan. Istilahnya kami cantumkan karena muncul di banyak nama, tanpa merekomendasikan model semacam itu. Model ini hanya ditujukan untuk penggunaan pribadi dan penelitian, dan tanggung jawab atas penggunaannya ada pada pengguna.