Przejdź do treści
fedi.software

Rodzaje modeli AI: mapa i słowniczek

LLM czy VLM, otwarte wagi czy open source, base, instruct, reasoning, destylowany, MoE, GGUF i kwantyzacja — co znaczą słowa w nazwach modeli, z przykładami na żywo.

Grupa Co oznacza Przykłady
LLMlarge language model · text model Rodzaj modelu A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal Rodzaj modelu An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights Otwartość The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition Otwartość The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model Otwartość No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base Wariant The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct Wariant The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking Wariant Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model Wariant Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen Wariant A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic Bez filtrów bezpieczeństwa Wariant A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model Architektura Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts Architektura The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active Architektura In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 Format pliku The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp Format pliku A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 Format pliku Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community Format pliku Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 Kwantyzacja GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL Kwantyzacja Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 Kwantyzacja Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

Modele oznaczone jako „bez filtrów bezpieczeństwa” (uncensored / abliterated) są wymienione wyłącznie do użytku osobistego i badawczego: nie mają ograniczeń bezpieczeństwa, a odpowiedzialność za ich użycie spoczywa na użytkowniku.

Zaktualizowano · Źródła: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

Osadź na swojej stronie

Wklej ten kod tam, gdzie ma się pojawić infografika: aktualizuje się sama. Użycie bezpłatne — zachowaj link do źródła.

JSON Nasze dane: CC BY 4.0 z linkiem do źródła; dane zewnętrzne zachowują licencję swojego źródła.

Mapa pojęć z nazw modeli

Nazwy otwartych modeli wyglądają jak szyfr, ale czyta się je po kawałku. Na początku stoi rodzina i całkowity rozmiar, w modelu MoE potem aktywne parametry (np. -A3B), dalej wariant (-Instruct lub -it), a na końcu format pliku i kwantyzacja pobieranego pliku. Słowniczek dzieli te pojęcia na sześć grup, a przy każdym terminie są aktualne przykłady – kliknięcie otwiera model w porównaniu modeli.

  • „Rodzaj modelu”: LLM pracuje z tekstem, VLM dodatkowo rozumie obrazy.
  • „Otwartość”: otwarte wagi, surowsza definicja Open Source AI Definition 1.0 od OSI albo modele zamknięte, dostępne tylko przez aplikację lub API.
  • „Wariant”: base, Instruct, reasoning, coder, destylowany oraz wersje społeczności z usuniętymi odmowami.
  • „Architektura”: model gęsty albo MoE, w którym na token pracuje tylko część parametrów.
  • „Format pliku”: Safetensors, GGUF, formaty wyłącznie dla GPU, jak AWQ, GPTQ i EXL, oraz MLX dla Maca.
  • „Kwantyzacja”: K-kwanty, jak Q4_K_M, pliki IQ i Unsloth Dynamic, FP8 i MXFP4.

Kolejność wyboru do lokalnego czatu

Weź wariant Instruct albo reasoning, nigdy model bazowy. Następnie sprawdź, ile masz pamięci, i dobierz kwant, który się zmieści: Q4_K_M to zwykły kompromis między rozmiarem a jakością, Q8_0 jest prawie bezstratny, a poniżej Q3 odpowiedzi wyraźnie słabną. Zostaw miejsce na potrzebny kontekst i dopiero na końcu oceń, czy szybkość ci odpowiada. Wykres który otwarty model pasuje do twojego sprzętu wykonuje te obliczenia dla znanych modeli.

Otwarte wagi to nie open source

Przy otwartych wagach (open weights) możesz pobrać plik, ale licencja może ograniczać użycie komercyjne albo wymagać akceptacji warunków, a dane treningowe zwykle pozostają zamknięte. To mniej, niż obiecuje pełne open source.

Destylacja to nie miniatura

Model destylowany jest mniejszym modelem wytrenowanym na odpowiedziach większego. Nie jest tym większym modelem w pomniejszeniu i wyraźnie mu ustępuje.

Abliterated i uncensored

To przeróbki otwartych modeli wykonane przez społeczność, w których usunięto odmawianie odpowiedzi. Wymieniamy ten termin, bo pojawia się w wielu nazwach, a nie dlatego, że takie modele polecamy. Są przeznaczone wyłącznie do użytku osobistego i badawczego, a odpowiedzialność za ich użycie spoczywa na użytkowniku.