Types de modèles d’IA : carte et glossaire
LLM ou VLM, poids ouverts ou open source, base, instruct, reasoning, distillé, MoE, GGUF et quantification — ce que signifient les mots dans les noms de modèles, avec des exemples en direct.
Type de modèle
Ouverture
Architecture
Format de fichier
Quantification
| Groupe | Ce que ça signifie | Exemples | |
|---|---|---|---|
| LLMlarge language model · text model | Type de modèle | A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it. | |
| VLMVL · vision-language · multimodal | Type de modèle | An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights. | |
| Open weightsopen model · downloadable weights | Ouverture | The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed. | |
| Open Source AIOSAID · OSI definition | Ouverture | The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. | — |
| Proprietary (API only)closed · API model | Ouverture | No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. | — |
| Basepretrained · -Base | Variante | The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. | — |
| Instruct-it · -Chat · -Instruct | Variante | The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat. | |
| ReasoningThinking · R1 · -Thinking | Variante | Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply. | |
| Coder-Coder · code model | Variante | Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model. | |
| Distill-Distill · R1-Distill-Qwen | Variante | A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. | — |
| Abliterated / uncensoreduncensored · abliterated · heretic Sans filtres de sécurité | Variante | A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. | — |
| Densedense model | Architecture | Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token. | |
| MoEmixture of experts | Architecture | The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM. | |
| Active parameters (-A3B)-A3B · -A22B · active | Architecture | In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number. | |
| Safetensors (BF16)safetensors · BF16 · FP16 | Format de fichier | The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. | — |
| GGUF.gguf · llama.cpp | Format de fichier | A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). | — |
| AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 | Format de fichier | Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. | — |
| MLXmlx-community | Format de fichier | Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. | — |
| Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 | Quantification | GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. | — |
| IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL | Quantification | Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable. | |
| FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 | Quantification | Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB. |
Les modèles marqués « sans filtres de sécurité » (uncensored / abliterated) ne sont listés qu’à des fins personnelles et de recherche : ils n’ont aucune restriction de sécurité, et la responsabilité de leur utilisation incombe à l’utilisateur.
Mis à jour · Sources: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)
Intégrer sur votre site
Collez ce code là où l’infographie doit apparaître : elle se met à jour toute seule. Utilisation gratuite — conservez le lien d’attribution.
Un nom de modèle se lit par morceaux
Les noms de modèles ressemblent à des codes, mais ils suivent presque toujours le même ordre. On trouve d'abord la famille et la taille, puis, pour un MoE, les paramètres actifs (-A3B), ensuite la variante (-Instruct ou -it). Le format de fichier et la quantification, eux, décrivent le téléchargement. Ce glossaire range ces termes en six groupes et relie chacun à des exemples vivants dans la comparaison de modèles.
- « Type de modèle » : LLM pour le texte, VLM pour le texte et l'image.
- « Ouverture » : poids ouverts (open weights), Open Source AI au sens de la définition 1.0 de l'OSI, modèles propriétaires.
- « Variante » : base, Instruct, reasoning, coder, distillé, abliterated / uncensored.
- « Architecture » : modèles denses, MoE, paramètres actifs.
- « Format de fichier » : Safetensors, GGUF, AWQ / GPTQ / EXL, MLX.
- « Quantification » : K-quants comme Q4_K_M, IQ et Unsloth Dynamic, FP8 / MXFP4.
Trois confusions fréquentes
Des poids ouverts ne font pas un modèle open source : vous pouvez télécharger les poids, mais leur licence peut restreindre l'usage et les données d'entraînement restent en général fermées. Un modèle distillé est un modèle plus petit entraîné sur les réponses d'un plus grand, pas le grand modèle lui-même. Enfin, côté quantification, Q4_K_M offre l'équilibre habituel entre taille et qualité, Q8_0 est presque sans perte, et sous Q3 la qualité baisse nettement.
Les variantes abliterated ou uncensored ont vu leur comportement de refus supprimé. Elles sont citées de façon neutre, pour un usage personnel et de recherche uniquement ; la responsabilité de leur utilisation incombe à l'utilisateur, et leur présence ne vaut pas recommandation.
Du glossaire à un chat local
Pour discuter avec un modèle en local, le raisonnement s'enchaîne ainsi : une variante Instruct (ou reasoning), puis la mémoire dont vous disposez, une quantification qui y tient, le contexte nécessaire et enfin la vitesse que vous acceptez. Les deux dernières étapes se vérifient dans quel modèle ouvert convient à votre matériel. Les définitions s'appuient sur les fiches et configurations des modèles sur Hugging Face, la documentation de llama.cpp et la définition Open Source AI de l'OSI.