Salt la conținut
fedi.software

Tipuri de modele de IA: hartă și glosar

LLM sau VLM, ponderi deschise sau open source, base, instruct, reasoning, distilat, MoE, GGUF și cuantizare — ce înseamnă cuvintele din numele modelelor, cu exemple în timp real.

Grup Ce înseamnă Exemple
LLMlarge language model · text model Tip de model A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal Tip de model An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights Deschidere The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition Deschidere The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model Deschidere No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base Variantă The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct Variantă The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking Variantă Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model Variantă Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen Variantă A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic Fără filtre de siguranță Variantă A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model Arhitectură Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts Arhitectură The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active Arhitectură In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 Format de fișier The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp Format de fișier A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 Format de fișier Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community Format de fișier Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 Cuantizare GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL Cuantizare Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 Cuantizare Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

Modelele marcate „fără filtre de siguranță” (uncensored / abliterated) sunt listate doar pentru uz personal și de cercetare: nu au restricții de siguranță, iar răspunderea pentru utilizare revine utilizatorului.

Actualizat · Surse: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

Încorporați pe site-ul dumneavoastră

Lipiți acest cod acolo unde trebuie să apară infograficul: se actualizează singur. Gratuit — păstrați linkul de atribuire.

JSON Datele noastre: CC BY 4.0 cu link către sursă; datele terților își păstrează licența sursei.

Denumirea unui model, bucată cu bucată

Numele modelelor par coduri, dar urmează aproape mereu aceeași ordine: familia și dimensiunea, apoi, la un MoE, parametrii activi (-A3B), după aceea varianta (-Instruct sau -it). Formatul fișierului și cuantizarea țin deja de descărcare. Glosarul împarte acești termeni în șase grupuri și leagă fiecare termen de exemple reale în comparația de modele.

Cele șase grupuri

  • „Tip de model”: LLM pentru text, VLM pentru text și imagine.
  • „Deschidere”: ponderi deschise (open weights), Open Source AI conform definiției 1.0 a OSI, modele proprietare.
  • „Variantă”: base, Instruct, reasoning, coder, distilat, abliterated / uncensored.
  • „Arhitectură”: modele dense, MoE, parametri activi.
  • „Format de fișier”: Safetensors, GGUF, AWQ / GPTQ / EXL, MLX.
  • „Cuantizare”: K-quants precum Q4_K_M, IQ și Unsloth Dynamic, FP8 / MXFP4.

Termeni care se confundă des

Ponderile deschise nu înseamnă open source. Ponderile se pot descărca, însă licența lor poate limita utilizarea, iar datele de antrenare rămân de obicei închise. Un model distilat este un model mai mic antrenat pe răspunsurile unuia mai mare, nu modelul mare în sine. La cuantizare, Q4_K_M este echilibrul obișnuit între dimensiune și calitate, Q8_0 aproape nu pierde nimic, iar sub Q3 calitatea scade vizibil.

Variantele abliterated sau uncensored au comportamentul de refuz eliminat. Le prezentăm neutru, doar pentru uz personal și de cercetare; responsabilitatea pentru utilizarea lor îi revine utilizatorului, iar prezența lor în listă nu este o recomandare.

De la glosar la un chat local

Pentru un chat pe propriul calculator, drumul arată așa: o variantă Instruct (sau reasoning), apoi memoria de care dispuneți, o cuantizare care încape în ea, contextul necesar și, la final, viteza pe care o acceptați. Ultimii pași îi verificați în ce model deschis încape pe hardware-ul dumneavoastră. Definițiile se bazează pe fișele și configurațiile modelelor de pe Hugging Face și pe documentația llama.cpp.