Soorten AI-modellen: een kaart en een woordenlijst
LLM of VLM, open weights of open source, base, instruct, reasoning, gedistilleerd, MoE, GGUF en quantisatie — wat de woorden in modelnamen betekenen, met live voorbeelden.
Modelsoort
Openheid
Architectuur
Bestandsformaat
Quantisatie
| Groep | Wat het betekent | Voorbeelden | |
|---|---|---|---|
| LLMlarge language model · text model | Modelsoort | A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it. | |
| VLMVL · vision-language · multimodal | Modelsoort | An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights. | |
| Open weightsopen model · downloadable weights | Openheid | The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed. | |
| Open Source AIOSAID · OSI definition | Openheid | The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. | — |
| Proprietary (API only)closed · API model | Openheid | No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. | — |
| Basepretrained · -Base | Variant | The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. | — |
| Instruct-it · -Chat · -Instruct | Variant | The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat. | |
| ReasoningThinking · R1 · -Thinking | Variant | Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply. | |
| Coder-Coder · code model | Variant | Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model. | |
| Distill-Distill · R1-Distill-Qwen | Variant | A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. | — |
| Abliterated / uncensoreduncensored · abliterated · heretic Geen veiligheidsfilters | Variant | A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. | — |
| Densedense model | Architectuur | Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token. | |
| MoEmixture of experts | Architectuur | The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM. | |
| Active parameters (-A3B)-A3B · -A22B · active | Architectuur | In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number. | |
| Safetensors (BF16)safetensors · BF16 · FP16 | Bestandsformaat | The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. | — |
| GGUF.gguf · llama.cpp | Bestandsformaat | A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). | — |
| AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 | Bestandsformaat | Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. | — |
| MLXmlx-community | Bestandsformaat | Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. | — |
| Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 | Quantisatie | GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. | — |
| IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL | Quantisatie | Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable. | |
| FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 | Quantisatie | Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB. |
Modellen met de aanduiding “geen veiligheidsfilters” (uncensored / abliterated) staan er alleen voor persoonlijk gebruik en onderzoek in: ze hebben geen veiligheidsbeperkingen en de verantwoordelijkheid voor het gebruik ligt bij de gebruiker.
Bijgewerkt · Bronnen: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)
Insluiten op je site
Plak deze code waar de infographic moet verschijnen: hij werkt zichzelf bij. Gratis te gebruiken — behoud de bronvermelding met link.
Waarom zijn modelnamen zo lang?
Omdat er een hele reeks keuzes in zit. Een naam lees je in stukjes: eerst de familie en de totale grootte, bij een MoE-model daarna de actieve parameters (zoals -A3B), dan de variant (-Instruct of -it) en ten slotte het bestandsformaat en de quantisatie van de download. De woordenlijst op deze pagina verdeelt die termen over zes groepen: „Modelsoort”, „Openheid”, „Variant”, „Architectuur”, „Bestandsformaat” en „Quantisatie”. Bij elke term staan live voorbeelden; een klik opent zo'n model in de modelvergelijking.
Wat betekenen de belangrijkste termen?
- LLM en VLM: een LLM werkt met tekst, een VLM kan daarnaast afbeeldingen lezen.
- Base en Instruct: een basismodel vult alleen tekst aan; een instructievariant (Instruct) is getraind om opdrachten op te volgen. Daarnaast bestaan reasoning-, coder- en gedistilleerde varianten.
- Dicht of MoE: een dicht model gebruikt al zijn parameters voor elk token, een mixture-of-experts-model (MoE) maar een deel.
- Bestanden: Safetensors is het gewone formaat, GGUF is bedoeld voor llama.cpp en de apps die daarop bouwen, AWQ, GPTQ en EXL draaien alleen op een GPU, en MLX is er voor Macs.
- Quants: K-quants zoals Q4_K_M, IQ- en Unsloth Dynamic-bestanden, FP8 en MXFP4.
Is open weights hetzelfde als open source?
Nee. Bij open gewichten (open weights) mag je het bestand downloaden, maar de licentie kan commercieel gebruik beperken of eisen dat je voorwaarden accepteert, en de trainingsdata blijft meestal gesloten. De OSI heeft met de Open Source AI Definition 1.0 een strengere norm opgesteld. Daarnaast zijn er gesloten modellen, die je alleen via een app of API gebruikt.
Is een gedistilleerd model een kleinere versie van het grote?
Niet echt. Het is een kleiner model dat getraind is op de antwoorden van een groter model. Het lijkt in stijl op het origineel, maar het is een ander en duidelijk zwakker model.
Welke variant neem ik voor een lokale chat?
Begin met een Instruct- of reasoning-variant, nooit met een basismodel. Kijk dan hoeveel geheugen je hebt en kies een quant die past: Q4_K_M is de gangbare balans tussen grootte en kwaliteit, Q8_0 is vrijwel verliesvrij en onder Q3 gaan de antwoorden merkbaar achteruit. Houd ruimte over voor de context die je nodig hebt en beoordeel pas daarna of de snelheid acceptabel is. De grafiek welk open model past op jouw hardware doet dat rekenwerk voor bekende modellen.
En abliterated of uncensored?
Dat zijn aanpassingen van open modellen door de community, waarbij het weigergedrag is verwijderd. We noemen de term omdat hij in veel modelnamen voorkomt, niet als aanbeveling. Zulke modellen zijn alleen bedoeld voor persoonlijk gebruik en onderzoek, en de verantwoordelijkheid voor het gebruik ligt bij de gebruiker.