{"key":"model-types","locale":"sl","title":"Vrste modelov umetne inteligence: zemljevid in slovar","url":"https://fedi.software/sl/ai/infographics/model-types","fields":[{"title":"Izraz","key":"term","format":"text","visible":true,"unit":null},{"title":"Skupina","key":"group","format":"badge","visible":true,"unit":null},{"title":"Kaj pomeni","key":"text","format":"text","visible":true,"unit":null},{"title":"Primeri","key":"examples","format":"list","visible":true,"unit":null}],"rows":[{"_id":"llm","term":"LLM","aka":["large language model","text model"],"group":"Vrsta modela","text":"A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.","examples":[{"id":"llama-3-3-70b-instruct","text":"Llama 3.3 70B Instruct","href":"https://fedi.software/sl/ai/compare?m=llama-3-3-70b-instruct","hf":"meta-llama/Llama-3.3-70B-Instruct"},{"id":"qwen3-32b","text":"Qwen3 32B","href":"https://fedi.software/sl/ai/compare?m=qwen3-32b","hf":"Qwen/Qwen3-32B"},{"id":"mistral-small-3-2-24b-instruct","text":"Mistral Small 3.2 24B","href":"https://fedi.software/sl/ai/compare?m=mistral-small-3-2-24b-instruct","hf":"mistralai/Mistral-Small-3.2-24B-Instruct-2506"}]},{"_id":"vlm","term":"VLM","aka":["VL","vision-language","multimodal"],"group":"Vrsta modela","text":"An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.","examples":[{"id":"qwen3-vl-30b-a3b-instruct","text":"Qwen3 VL 30B A3B Instruct","href":"https://fedi.software/sl/ai/compare?m=qwen3-vl-30b-a3b-instruct","hf":"Qwen/Qwen3-VL-30B-A3B-Instruct"},{"id":"gemma-3-27b-it","text":"Gemma 3 27B","href":"https://fedi.software/sl/ai/compare?m=gemma-3-27b-it","hf":"google/gemma-3-27b-it"},{"id":"qwen2-5-vl-7b-instruct","text":"Qwen2.5-VL 7B Instruct","href":"https://fedi.software/sl/ai/compare?m=qwen2-5-vl-7b-instruct"}]},{"_id":"open-weights","term":"Open weights","aka":["open model","downloadable weights"],"group":"Odprtost","text":"The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.","examples":[{"id":"llama-3-3-70b-instruct","text":"Llama 3.3 70B Instruct","href":"https://fedi.software/sl/ai/compare?m=llama-3-3-70b-instruct","hf":"meta-llama/Llama-3.3-70B-Instruct"},{"id":"gemma-3-27b-it","text":"Gemma 3 27B","href":"https://fedi.software/sl/ai/compare?m=gemma-3-27b-it","hf":"google/gemma-3-27b-it"},{"id":"openai-gpt-oss-120b","text":"gpt-oss-120b","href":"https://fedi.software/sl/ai/compare?m=openai-gpt-oss-120b","hf":"openai/gpt-oss-120b"}]},{"_id":"open-source-ai","term":"Open Source AI","aka":["OSAID","OSI definition"],"group":"Odprtost","text":"The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source.","examples":[]},{"_id":"proprietary","term":"Proprietary (API only)","aka":["closed","API model"],"group":"Odprtost","text":"No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally.","examples":[]},{"_id":"base","term":"Base","aka":["pretrained","-Base"],"group":"Različica","text":"The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting.","examples":[]},{"_id":"instruct","term":"Instruct","aka":["-it","-Chat","-Instruct"],"group":"Različica","text":"The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.","examples":[{"id":"gemma-3-27b-it","text":"Gemma 3 27B","href":"https://fedi.software/sl/ai/compare?m=gemma-3-27b-it","hf":"google/gemma-3-27b-it"},{"id":"llama-3-1-8b-instruct","text":"Llama 3.1 8B Instruct","href":"https://fedi.software/sl/ai/compare?m=llama-3-1-8b-instruct","hf":"meta-llama/Meta-Llama-3.1-8B-Instruct"},{"id":"qwen3-30b-a3b-instruct-2507","text":"Qwen3 30B A3B Instruct 2507","href":"https://fedi.software/sl/ai/compare?m=qwen3-30b-a3b-instruct-2507","hf":"Qwen/Qwen3-30B-A3B-Instruct-2507"}]},{"_id":"reasoning","term":"Reasoning","aka":["Thinking","R1","-Thinking"],"group":"Različica","text":"Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.","examples":[{"id":"deepseek-r1","text":"DeepSeek R1","href":"https://fedi.software/sl/ai/compare?m=deepseek-r1","hf":"deepseek-ai/DeepSeek-R1"},{"id":"qwen3-30b-a3b-thinking-2507","text":"Qwen3 30B A3B Thinking 2507","href":"https://fedi.software/sl/ai/compare?m=qwen3-30b-a3b-thinking-2507","hf":"Qwen/Qwen3-30B-A3B-Thinking-2507"},{"id":"kimi-k2-thinking","text":"Kimi K2 Thinking","href":"https://fedi.software/sl/ai/compare?m=kimi-k2-thinking","hf":"moonshotai/Kimi-K2-Thinking"}]},{"_id":"coder","term":"Coder","aka":["-Coder","code model"],"group":"Različica","text":"Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.","examples":[{"id":"qwen3-coder-30b-a3b-instruct","text":"Qwen3 Coder 30B A3B Instruct","href":"https://fedi.software/sl/ai/compare?m=qwen3-coder-30b-a3b-instruct","hf":"Qwen/Qwen3-Coder-30B-A3B-Instruct"},{"id":"qwen2-5-coder-32b-instruct","text":"Qwen2.5 Coder 32B Instruct","href":"https://fedi.software/sl/ai/compare?m=qwen2-5-coder-32b-instruct","hf":"Qwen/Qwen2.5-Coder-32B-Instruct"},{"id":"deepseek-coder-v2","text":"deepseek-coder-v2","href":"https://fedi.software/sl/ai/compare?m=deepseek-coder-v2"}]},{"_id":"distill","term":"Distill","aka":["-Distill","R1-Distill-Qwen"],"group":"Različica","text":"A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original.","examples":[]},{"_id":"abliterated","_risk":true,"term":"Abliterated / uncensored","aka":["uncensored","abliterated","heretic"],"group":"Različica","text":"A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it.","examples":[]},{"_id":"dense","term":"Dense","aka":["dense model"],"group":"Arhitektura","text":"Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.","examples":[{"id":"llama-3-3-70b-instruct","text":"Llama 3.3 70B Instruct","href":"https://fedi.software/sl/ai/compare?m=llama-3-3-70b-instruct","hf":"meta-llama/Llama-3.3-70B-Instruct"},{"id":"gemma-3-27b-it","text":"Gemma 3 27B","href":"https://fedi.software/sl/ai/compare?m=gemma-3-27b-it","hf":"google/gemma-3-27b-it"},{"id":"qwen3-32b","text":"Qwen3 32B","href":"https://fedi.software/sl/ai/compare?m=qwen3-32b","hf":"Qwen/Qwen3-32B"}]},{"_id":"moe","term":"MoE","aka":["mixture of experts"],"group":"Arhitektura","text":"The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.","examples":[{"id":"openai-gpt-oss-120b","text":"gpt-oss-120b","href":"https://fedi.software/sl/ai/compare?m=openai-gpt-oss-120b","hf":"openai/gpt-oss-120b"},{"id":"qwen3-30b-a3b","text":"Qwen3 30B A3B","href":"https://fedi.software/sl/ai/compare?m=qwen3-30b-a3b","hf":"Qwen/Qwen3-30B-A3B"},{"id":"deepseek-r1","text":"DeepSeek R1","href":"https://fedi.software/sl/ai/compare?m=deepseek-r1","hf":"deepseek-ai/DeepSeek-R1"}]},{"_id":"active-params","term":"Active parameters (-A3B)","aka":["-A3B","-A22B","active"],"group":"Arhitektura","text":"In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.","examples":[{"id":"qwen3-30b-a3b","text":"Qwen3 30B A3B","href":"https://fedi.software/sl/ai/compare?m=qwen3-30b-a3b","hf":"Qwen/Qwen3-30B-A3B"},{"id":"qwen3-235b-a22b","text":"Qwen3 235B A22B","href":"https://fedi.software/sl/ai/compare?m=qwen3-235b-a22b","hf":"Qwen/Qwen3-235B-A22B"}]},{"_id":"safetensors","term":"Safetensors (BF16)","aka":["safetensors","BF16","FP16"],"group":"Format datoteke","text":"The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation.","examples":[]},{"_id":"gguf","term":"GGUF","aka":[".gguf","llama.cpp"],"group":"Format datoteke","text":"A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003).","examples":[]},{"_id":"gpu-formats","term":"AWQ / GPTQ / EXL2-3","aka":["AWQ","GPTQ","EXL2","EXL3"],"group":"Format datoteke","text":"Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM.","examples":[]},{"_id":"mlx","term":"MLX","aka":["mlx-community"],"group":"Format datoteke","text":"Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size.","examples":[]},{"_id":"k-quants","term":"Q4_K_M and other K-quants","aka":["Q4_K_M","Q5_K_M","Q6_K","Q8_0"],"group":"Kvantizacija","text":"GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably.","examples":[]},{"_id":"i-quants","term":"IQ quants and Unsloth Dynamic (UD)","aka":["IQ2_XXS","IQ3_K","UD-Q2_K_XL"],"group":"Kvantizacija","text":"Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.","examples":[{"id":"deepseek-r1","text":"DeepSeek R1","href":"https://fedi.software/sl/ai/compare?m=deepseek-r1","hf":"deepseek-ai/DeepSeek-R1"},{"id":"qwen3-235b-a22b","text":"Qwen3 235B A22B","href":"https://fedi.software/sl/ai/compare?m=qwen3-235b-a22b","hf":"Qwen/Qwen3-235B-A22B"}]},{"_id":"fp-formats","term":"FP8 / MXFP4 / NVFP4","aka":["FP8","MXFP4","NVFP4"],"group":"Kvantizacija","text":"Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.","examples":[{"id":"openai-gpt-oss-120b","text":"gpt-oss-120b","href":"https://fedi.software/sl/ai/compare?m=openai-gpt-oss-120b","hf":"openai/gpt-oss-120b"},{"id":"gpt-oss-20b","text":"gpt-oss-20b","href":"https://fedi.software/sl/ai/compare?m=gpt-oss-20b","hf":"openai/gpt-oss-20b"}]}],"meta":{"sources":[{"name":"fedi.software","url":"https://fedi.software/","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","modified":false},{"name":"Open Source Initiative (OSAID 1.0)","url":"https://opensource.org/ai/open-source-ai-definition","license":null,"license_url":null,"modified":false}],"updated":"2026-10-02","notes":[],"license":{"name":"fedi.software","license":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/","scope":"fedi.software data: hardware, builds, glossary, Fit calculations, curation and the compilation; third-party data keeps the licence of its source (sources[])"},"attribution":["Data by fedi.software, CC BY 4.0 — a link to the source page is required."]}}