Types of AI models: a map and a glossary
LLM or VLM, open weights or open source, base, instruct, reasoning, distilled, MoE, GGUF and quantization — what the words in model names mean, with live examples.
Model type
Openness
Architecture
File format
Quantization
| Group | What it means | Examples | |
|---|---|---|---|
| LLMlarge language model · text model | Model type | A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it. | |
| VLMVL · vision-language · multimodal | Model type | An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights. | |
| Open weightsopen model · downloadable weights | Openness | The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed. | |
| Open Source AIOSAID · OSI definition | Openness | The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. | — |
| Proprietary (API only)closed · API model | Openness | No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. | — |
| Basepretrained · -Base | Variant | The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. | — |
| Instruct-it · -Chat · -Instruct | Variant | The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat. | |
| ReasoningThinking · R1 · -Thinking | Variant | Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply. | |
| Coder-Coder · code model | Variant | Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model. | |
| Distill-Distill · R1-Distill-Qwen | Variant | A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. | — |
| Abliterated / uncensoreduncensored · abliterated · heretic No safety filters | Variant | A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. | — |
| Densedense model | Architecture | Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token. | |
| MoEmixture of experts | Architecture | The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM. | |
| Active parameters (-A3B)-A3B · -A22B · active | Architecture | In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number. | |
| Safetensors (BF16)safetensors · BF16 · FP16 | File format | The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. | — |
| GGUF.gguf · llama.cpp | File format | A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). | — |
| AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 | File format | Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. | — |
| MLXmlx-community | File format | Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. | — |
| Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 | Quantization | GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. | — |
| IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL | Quantization | Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable. | |
| FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 | Quantization | Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB. |
Models marked “no safety filters” (uncensored / abliterated) are listed for personal and research use only: they have no safety restrictions, and the responsibility for their use lies with the user.
Updated · Sources: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)
Embed on your site
Paste this code where the infographic should appear: it updates by itself. Free to use — keep the attribution link.
Reading a model name piece by piece
A name like "Qwen3-30B-A3B-Instruct-GGUF Q4_K_M" packs five decisions into one line. First comes the family and the total size; then, for a MoE model, the active parameters; then the variant; and finally the file format and quantization of the download. The glossary above splits these words into six groups, with live examples next to each term — a click on an example opens it in the model comparison.
The six groups in short
- Model type: an LLM works with text; a VLM also takes images.
- Openness: open weights, the stricter Open Source AI definition of the OSI, or proprietary models available only through an app or API.
- Variant: base, Instruct, reasoning, coder, distilled, and community builds with the refusals removed.
- Architecture: dense versus MoE, and the active-parameter suffix of MoE names.
- File format: Safetensors, GGUF, GPU-only formats such as AWQ, GPTQ and EXL, and MLX for Macs.
- Quantization: K-quants like Q4_K_M, IQ and Unsloth Dynamic files, FP8 and MXFP4.
From terms to a choice
For a local chat assistant the order of decisions is fairly fixed. Take an Instruct or reasoning variant, never a base model. Look at the memory you have, then pick a quant that fits in it: Q4_K_M is the usual balance of size and quality, Q8_0 is almost lossless, and below Q3 the answers degrade noticeably. Leave room for the context you need, and only then judge whether the speed is acceptable. The hardware fit chart does this arithmetic for well-known models.
Two common confusions
Open weights is not the same as open source. You can download the file, but the licence may restrict commercial use or ask you to accept terms, and the training data usually stays closed. A distilled model is also often misread: it is a smaller model trained on the answers of a bigger one, not the bigger model itself, and it is much weaker.
About abliterated and uncensored models
These are community modifications of open models with the refusal behaviour removed. We list the term because it appears in many names, without recommending such models. They are meant for personal and research use only, and the responsibility for how they are used lies with the user.