Na vsebino
fedi.software

Vrste modelov umetne inteligence: zemljevid in slovar

LLM ali VLM, odprte uteži ali odprta koda, base, instruct, reasoning, destilirani, MoE, GGUF in kvantizacija — kaj pomenijo besede v imenih modelov, z živimi primeri.

Skupina Kaj pomeni Primeri
LLMlarge language model · text model Vrsta modela A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal Vrsta modela An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights Odprtost The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition Odprtost The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model Odprtost No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base Različica The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct Različica The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking Različica Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model Različica Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen Različica A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic Brez varnostnih filtrov Različica A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model Arhitektura Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts Arhitektura The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active Arhitektura In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 Format datoteke The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp Format datoteke A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 Format datoteke Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community Format datoteke Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 Kvantizacija GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL Kvantizacija Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 Kvantizacija Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

Modeli, označeni z „brez varnostnih filtrov“ (uncensored / abliterated), so navedeni samo za osebno uporabo in raziskave: nimajo varnostnih omejitev, odgovornost za uporabo pa je na uporabniku.

Posodobljeno · Viri: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

Vdelaj na svojo stran

Prilepite to kodo tja, kjer naj se prikaže infografika: posodablja se sama. Brezplačna uporaba — ohranite povezavo do vira.

JSON Naši podatki: CC BY 4.0 s povezavo do vira; podatki tretjih oseb ohranijo licenco svojega vira.

Kako prebrati ime modela?

Ime, kot je »Qwen3-30B-A3B-Instruct-GGUF Q4_K_M«, v eni vrstici združi več odločitev. Najprej sta družina in skupna velikost, pri modelu MoE sledijo aktivni parametri (-A3B), nato različica (-Instruct ali -it), na koncu pa format datoteke in kvantizacija prenosa. Slovar na tej strani te besede razvrsti v šest skupin in ob vsakem izrazu pokaže žive primere; klik na primer ga odpre v primerjavi modelov.

Katere so te skupine?

  • Vrsta modela: LLM dela z besedilom, VLM sprejme tudi slike.
  • Odprtost: odprte uteži, strožja definicija Open Source AI organizacije OSI ali lastniški modeli, dostopni samo prek aplikacije ali API.
  • Različica: base, Instruct, reasoning, coder, destilirani modeli in skupnostne predelave brez zavračanja.
  • Arhitektura: gosti modeli proti MoE ter pripona z aktivnimi parametri.
  • Format datoteke: Safetensors, GGUF, formati samo za grafične procesorje (AWQ, GPTQ, EXL) in MLX za računalnike Mac.
  • Kvantizacija: K-kvanti, kot je Q4_K_M, datoteke IQ in Unsloth Dynamic, FP8 in MXFP4.

V kakšnem vrstnem redu izbirati za lokalni klepet?

Vzemite različico Instruct ali reasoning, osnovnega modela base pa ne. Nato poglejte, koliko pomnilnika imate, in izberite kvantizacijo, ki gre vanj. Q4_K_M je običajno ravnovesje med velikostjo in kakovostjo, Q8_0 je skoraj brez izgub, pod Q3 pa kakovost odgovorov opazno pade. Pustite prostor za kontekst, ki ga potrebujete, in šele potem presodite, ali je hitrost sprejemljiva. Ta račun za znane modele naredi infografika o strojni opremi.

Ali so odprte uteži isto kot odprta koda?

Niso. Datoteko z odprtimi utežmi lahko prenesete, licenca pa lahko omeji komercialno rabo ali zahteva sprejetje pogojev, učni podatki pa običajno ostanejo zaprti. Podobno pogosto narobe razumemo destilirani model: to je manjši model, naučen na odgovorih večjega, ne pa večji model sam, in je občutno šibkejši.

Kaj pomenita abliterated in uncensored?

Gre za skupnostne predelave odprtih modelov, ki jim je odstranjeno zavračanje zahtev. Izraz navajamo, ker se pojavlja v mnogih imenih, takih modelov pa ne priporočamo. Namenjeni so samo osebni uporabi in raziskavam, odgovornost za njihovo uporabo pa nosi uporabnik.