Τύποι μοντέλων ΤΝ: χάρτης και γλωσσάρι
LLM ή VLM, ανοιχτά βάρη ή ανοιχτός κώδικας, base, instruct, reasoning, distilled, MoE, GGUF και κβαντοποίηση — τι σημαίνουν οι όροι στα ονόματα των μοντέλων, με ζωντανά παραδείγματα.
Τύπος μοντέλου
Ανοιχτότητα
Αρχιτεκτονική
Μορφή αρχείου
Κβαντοποίηση
| Ομάδα | Τι σημαίνει | Παραδείγματα | |
|---|---|---|---|
| LLMlarge language model · text model | Τύπος μοντέλου | A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it. | |
| VLMVL · vision-language · multimodal | Τύπος μοντέλου | An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights. | |
| Open weightsopen model · downloadable weights | Ανοιχτότητα | The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed. | |
| Open Source AIOSAID · OSI definition | Ανοιχτότητα | The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. | — |
| Proprietary (API only)closed · API model | Ανοιχτότητα | No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. | — |
| Basepretrained · -Base | Παραλλαγή | The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. | — |
| Instruct-it · -Chat · -Instruct | Παραλλαγή | The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat. | |
| ReasoningThinking · R1 · -Thinking | Παραλλαγή | Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply. | |
| Coder-Coder · code model | Παραλλαγή | Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model. | |
| Distill-Distill · R1-Distill-Qwen | Παραλλαγή | A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. | — |
| Abliterated / uncensoreduncensored · abliterated · heretic Χωρίς φίλτρα ασφαλείας | Παραλλαγή | A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. | — |
| Densedense model | Αρχιτεκτονική | Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token. | |
| MoEmixture of experts | Αρχιτεκτονική | The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM. | |
| Active parameters (-A3B)-A3B · -A22B · active | Αρχιτεκτονική | In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number. | |
| Safetensors (BF16)safetensors · BF16 · FP16 | Μορφή αρχείου | The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. | — |
| GGUF.gguf · llama.cpp | Μορφή αρχείου | A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). | — |
| AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 | Μορφή αρχείου | Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. | — |
| MLXmlx-community | Μορφή αρχείου | Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. | — |
| Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 | Κβαντοποίηση | GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. | — |
| IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL | Κβαντοποίηση | Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable. | |
| FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 | Κβαντοποίηση | Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB. |
Τα μοντέλα με την ένδειξη «Χωρίς φίλτρα ασφαλείας» (uncensored / abliterated) παρατίθενται μόνο για προσωπική και ερευνητική χρήση: δεν έχουν περιορισμούς ασφαλείας και την ευθύνη για τη χρήση τους φέρει ο χρήστης.
Ενημερώθηκε · Πηγές: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)
Ενσωμάτωση στον ιστότοπό σας
Επικολλήστε αυτόν τον κώδικα εκεί που πρέπει να εμφανίζεται το infographic: ενημερώνεται μόνο του. Δωρεάν χρήση — κρατήστε τον σύνδεσμο αναφοράς.
Ένα όνομα, πέντε αποφάσεις
Ένα όνομα όπως «Qwen3-30B-A3B-Instruct-GGUF Q4_K_M» συμπυκνώνει πολλές πληροφορίες σε μία γραμμή. Πρώτα έρχονται η οικογένεια και το συνολικό μέγεθος, σε ένα μοντέλο MoE ακολουθούν οι ενεργές παράμετροι (-A3B), μετά η παραλλαγή (-Instruct ή -it) και στο τέλος η μορφή αρχείου και η κβαντοποίηση της λήψης. Το γλωσσάρι της σελίδας χωρίζει αυτές τις λέξεις σε έξι ομάδες, με ζωντανά παραδείγματα δίπλα σε κάθε όρο· ένα κλικ σε παράδειγμα το ανοίγει στη σύγκριση μοντέλων.
Οι έξι ομάδες
- Τύπος μοντέλου: ένα LLM δουλεύει με κείμενο, ένα VLM δέχεται και εικόνες.
- Ανοιχτότητα: ανοιχτά βάρη, ο αυστηρότερος ορισμός Open Source AI του OSI ή ιδιόκτητα μοντέλα, διαθέσιμα μόνο μέσω εφαρμογής ή API.
- Παραλλαγή: base, Instruct, reasoning, coder, distilled και κοινοτικές εκδοχές χωρίς αρνήσεις.
- Αρχιτεκτονική: πυκνά μοντέλα έναντι MoE και η κατάληξη των ενεργών παραμέτρων.
- Μορφή αρχείου: Safetensors, GGUF, μορφές μόνο για GPU (AWQ, GPTQ, EXL) και MLX για Mac.
- Κβαντοποίηση: K-quants όπως το Q4_K_M, αρχεία IQ και Unsloth Dynamic, FP8 και MXFP4.
Από τους όρους στην επιλογή, βήμα βήμα
- Για τοπική συνομιλία πάρτε παραλλαγή Instruct ή reasoning, όχι ένα μοντέλο base.
- Δείτε πόση μνήμη διαθέτετε.
- Διαλέξτε κβαντοποίηση που χωράει: το Q4_K_M είναι η συνήθης ισορροπία μεγέθους και ποιότητας, το Q8_0 σχεδόν χωρίς απώλειες, ενώ κάτω από Q3 οι απαντήσεις χειροτερεύουν αισθητά.
- Αφήστε χώρο για το context που χρειάζεστε.
- Μόνο τότε κρίνετε αν η ταχύτητα σας ικανοποιεί — τον λογαριασμό για γνωστά μοντέλα τον κάνει το γράφημα συμβατότητας υλικού.
Δύο συνηθισμένες παρανοήσεις
Τα ανοιχτά βάρη δεν ταυτίζονται με τον ανοιχτό κώδικα. Το αρχείο κατεβαίνει ελεύθερα, όμως η άδεια μπορεί να περιορίζει την εμπορική χρήση ή να ζητά αποδοχή όρων, και τα δεδομένα εκπαίδευσης συνήθως μένουν κλειστά. Επίσης, ένα distilled μοντέλο είναι μικρότερο μοντέλο εκπαιδευμένο στις απαντήσεις ενός μεγαλύτερου, όχι το ίδιο το μεγάλο μοντέλο, και είναι αισθητά πιο αδύναμο.
Abliterated και uncensored
Έτσι λέγονται κοινοτικές τροποποιήσεις ανοιχτών μοντέλων από τις οποίες έχει αφαιρεθεί η άρνηση αιτημάτων. Αναφέρουμε τον όρο επειδή εμφανίζεται σε πολλά ονόματα, χωρίς να προτείνουμε τέτοια μοντέλα. Προορίζονται αποκλειστικά για προσωπική χρήση και έρευνα, και την ευθύνη για τη χρήση τους φέρει ο χρήστης.