Tipi di modelli di IA: mappa e glossario
LLM o VLM, pesi aperti o open source, base, instruct, reasoning, distillato, MoE, GGUF e quantizzazione — cosa significano le parole nei nomi dei modelli, con esempi dal vivo.
Tipo di modello
Apertura
Architettura
Formato del file
Quantizzazione
| Gruppo | Cosa significa | Esempi | |
|---|---|---|---|
| LLMlarge language model · text model | Tipo di modello | A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it. | |
| VLMVL · vision-language · multimodal | Tipo di modello | An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights. | |
| Open weightsopen model · downloadable weights | Apertura | The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed. | |
| Open Source AIOSAID · OSI definition | Apertura | The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. | — |
| Proprietary (API only)closed · API model | Apertura | No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. | — |
| Basepretrained · -Base | Variante | The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. | — |
| Instruct-it · -Chat · -Instruct | Variante | The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat. | |
| ReasoningThinking · R1 · -Thinking | Variante | Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply. | |
| Coder-Coder · code model | Variante | Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model. | |
| Distill-Distill · R1-Distill-Qwen | Variante | A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. | — |
| Abliterated / uncensoreduncensored · abliterated · heretic Senza filtri di sicurezza | Variante | A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. | — |
| Densedense model | Architettura | Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token. | |
| MoEmixture of experts | Architettura | The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM. | |
| Active parameters (-A3B)-A3B · -A22B · active | Architettura | In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number. | |
| Safetensors (BF16)safetensors · BF16 · FP16 | Formato del file | The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. | — |
| GGUF.gguf · llama.cpp | Formato del file | A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). | — |
| AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 | Formato del file | Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. | — |
| MLXmlx-community | Formato del file | Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. | — |
| Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 | Quantizzazione | GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. | — |
| IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL | Quantizzazione | Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable. | |
| FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 | Quantizzazione | Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB. |
I modelli contrassegnati come «senza filtri di sicurezza» (uncensored / abliterated) sono elencati solo per uso personale e di ricerca: non hanno restrizioni di sicurezza e la responsabilità del loro utilizzo ricade sull’utente.
Aggiornato · Fonti: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)
Incorpora nel tuo sito
Incolla questo codice dove deve apparire l’infografica: si aggiorna da sola. Uso gratuito — mantieni il link di attribuzione.
Come si legge il nome di un modello?
A pezzi, quasi sempre nello stesso ordine: famiglia e dimensione, poi per un MoE i parametri attivi (-A3B), quindi la variante (-Instruct o -it). Formato del file e quantizzazione riguardano invece il download. Il glossario divide questi termini in sei gruppi e collega ciascuno a esempi reali nel confronto tra modelli.
Quali sono i sei gruppi?
- «Tipo di modello»: LLM per il testo, VLM per testo e immagini.
- «Apertura»: pesi aperti (open weights), Open Source AI secondo la definizione 1.0 della OSI, modelli proprietari.
- «Variante»: base, Instruct, reasoning, coder, distillato, abliterated / uncensored.
- «Architettura»: modelli densi, MoE, parametri attivi.
- «Formato del file»: Safetensors, GGUF, AWQ / GPTQ / EXL, MLX.
- «Quantizzazione»: K-quant come Q4_K_M, IQ e Unsloth Dynamic, FP8 / MXFP4.
Pesi aperti e open source sono la stessa cosa?
No. I pesi si possono scaricare, ma la loro licenza può limitarne l'uso e i dati di addestramento di solito restano chiusi. Allo stesso modo, un modello distillato è un modello più piccolo addestrato sulle risposte di uno più grande, non il modello grande in versione ridotta.
Quale quantizzazione scegliere?
Q4_K_M è il compromesso abituale tra dimensione e qualità, Q8_0 è quasi senza perdite, sotto Q3 la qualità cala in modo evidente.
Cosa sono le varianti abliterated o uncensored?
Sono modelli a cui è stato rimosso il comportamento di rifiuto. Li elenchiamo in modo neutrale, solo per uso personale e di ricerca; la responsabilità del loro utilizzo è dell'utente e la loro presenza non è una raccomandazione.
Come arrivo a una chat locale?
Scegli una variante Instruct (o reasoning), guarda quanta memoria hai, trova una quantizzazione che ci stia, decidi il contesto che ti serve e infine la velocità che ti va bene. Gli ultimi passaggi li verifichi in quale modello aperto entra nel tuo hardware. Le definizioni si basano sulle schede e sulle configurazioni dei modelli su Hugging Face e sulla documentazione di llama.cpp.