Zum Inhalt springen
fedi.software

Arten von KI-Modellen: Übersicht und Glossar

LLM oder VLM, Open Weights oder Open Source, Base, Instruct, Reasoning, destilliert, MoE, GGUF und Quantisierung — was die Begriffe in Modellnamen bedeuten, mit Live-Beispielen.

Gruppe Bedeutung Beispiele
LLMlarge language model · text model Modelltyp A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal Modelltyp An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights Offenheit The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition Offenheit The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model Offenheit No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base Variante The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct Variante The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking Variante Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model Variante Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen Variante A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic Ohne Sicherheitsfilter Variante A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model Architektur Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts Architektur The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active Architektur In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 Dateiformat The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp Dateiformat A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 Dateiformat Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community Dateiformat Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 Quantisierung GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL Quantisierung Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 Quantisierung Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

Modelle mit der Kennzeichnung «Ohne Sicherheitsfilter» (uncensored / abliterated) sind nur für den persönlichen und den Forschungsgebrauch aufgeführt: Sie haben keine Sicherheitsbeschränkungen, und die Verantwortung für ihre Nutzung liegt beim Nutzer.

Aktualisiert · Quellen: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

Auf deiner Website einbetten

Füge diesen Code dort ein, wo die Infografik erscheinen soll: Sie aktualisiert sich von selbst. Kostenlos nutzbar — lass den Quellenlink stehen.

JSON Unsere Daten: CC BY 4.0 mit Link zur Quelle; Daten Dritter behalten die Lizenz ihrer Quelle.

Ein Modellname in Einzelteilen

Lange Modellnamen mit Bindestrichen wirken kryptisch, folgen aber einer festen Reihenfolge. Zuerst kommen Familie und Grösse, bei MoE-Modellen dann die aktiven Parameter (-A3B), danach die Variante (-Instruct oder -it). Dateiformat und Quantisierung beziehen sich bereits auf den Download. Das Glossar ordnet alle diese Begriffe sechs Gruppen zu, und zu jedem Begriff verlinkt es Live-Beispiele im Modellvergleich.

Die sechs Gruppen im Überblick

  • «Modelltyp»: LLM für Text, VLM für Text und Bild.
  • «Offenheit»: offene Gewichte (open weights), Open Source AI nach der Definition 1.0 der OSI und proprietäre Modelle.
  • «Variante»: Base, Instruct, Reasoning, Coder, destilliert sowie abliterated / uncensored.
  • «Architektur»: dichte Modelle, MoE und aktive Parameter.
  • «Dateiformat»: Safetensors, GGUF, AWQ / GPTQ / EXL und MLX.
  • «Quantisierung»: K-Quants wie Q4_K_M, IQ und Unsloth Dynamic, FP8 / MXFP4.

Begriffe, die oft verwechselt werden

Offene Gewichte sind nicht dasselbe wie Open Source. Die Gewichte lassen sich herunterladen, können aber unter einer Lizenz stehen, die die Nutzung einschränkt, und die Trainingsdaten bleiben meist verschlossen. Ein destilliertes Modell ist ein kleineres Modell, das an den Antworten eines grösseren gelernt hat, nicht das grössere Modell selbst. Bei der Quantisierung gilt Q4_K_M als übliche Balance aus Dateigrösse und Qualität, Q8_0 als nahezu verlustfrei, und unterhalb von Q3 sinkt die Qualität merklich.

Varianten mit der Kennzeichnung abliterated oder uncensored haben das Ablehnungsverhalten entfernt. Wir führen sie neutral und nur für den persönlichen und den Forschungsgebrauch auf; die Verantwortung für die Nutzung liegt beim Nutzer, eine Empfehlung ist damit nicht verbunden.

Vom Begriff zum lokalen Chat

Für einen lokalen Chat führt der Weg über eine Instruct- oder Reasoning-Variante, dann über den vorhandenen Speicher zu einer passenden Quantisierung, anschliessend zum nötigen Kontext und zuletzt zur Geschwindigkeit, mit der Sie leben können. Speicher und Tempo rechnet Welches offene Modell zu Ihrer Hardware passt für Sie aus. Grundlage des Glossars sind Model Cards und Konfigurationen auf Hugging Face sowie die Dokumentation von llama.cpp.