İçeriğe geç
fedi.software

Yapay zekâ modeli türleri: harita ve sözlük

LLM veya VLM, açık ağırlık veya açık kaynak, base, instruct, reasoning, damıtılmış, MoE, GGUF ve nicemleme — model adlarındaki sözcüklerin anlamı, canlı örneklerle.

Grup Anlamı Örnekler
LLMlarge language model · text model Model türü A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal Model türü An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights Açıklık The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition Açıklık The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model Açıklık No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base Varyant The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct Varyant The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking Varyant Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model Varyant Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen Varyant A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic Güvenlik filtresi yok Varyant A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model Mimari Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts Mimari The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active Mimari In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 Dosya biçimi The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp Dosya biçimi A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 Dosya biçimi Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community Dosya biçimi Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 Nicemleme GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL Nicemleme Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 Nicemleme Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

“Güvenlik filtresi yok” işaretli modeller (uncensored / abliterated) yalnızca kişisel ve araştırma amaçlı listelenmiştir: güvenlik kısıtlamaları yoktur ve kullanım sorumluluğu kullanıcıya aittir.

Güncellendi · Kaynaklar: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

Sitenize ekleyin

Bu kodu infografiğin görünmesini istediğiniz yere yapıştırın: kendiliğinden güncellenir. Ücretsiz kullanım — kaynak bağlantısını koruyun.

JSON Verilerimiz: kaynağa bağlantıyla CC BY 4.0; üçüncü taraf verileri kaynağının lisansını korur.

Bir model adı nasıl çözülür

«Qwen3-30B-A3B-Instruct-GGUF Q4_K_M» gibi bir ad, tek satıra birkaç kararı sığdırır. Önce aile ve toplam boyut gelir; MoE modelinde ardından etkin parametreler (-A3B), sonra varyant (-Instruct ya da -it), en sonda da indirilen dosyanın biçimi ve nicemlemesi. Sayfadaki sözlük bu kelimeleri altı gruba ayırır ve her terimin yanında canlı örnekler gösterir; bir örneğe tıklamak onu model karşılaştırmasında açar.

Altı grup kısaca

  • Model türü: LLM metinle çalışır, VLM görselleri de alır.
  • Açıklık: açık ağırlıklar, OSI'nin daha katı Open Source AI tanımı ya da yalnızca uygulama veya API üzerinden erişilen tescilli modeller.
  • Varyant: base, Instruct, reasoning, coder, damıtılmış (distilled) modeller ve reddetme davranışı kaldırılmış topluluk sürümleri.
  • Mimari: yoğun modeller ile MoE ve etkin parametre eki.
  • Dosya biçimi: Safetensors, GGUF, yalnızca GPU için olan AWQ, GPTQ ve EXL, Mac'ler için MLX.
  • Nicemleme: Q4_K_M gibi K-quant'lar, IQ ve Unsloth Dynamic dosyaları, FP8 ve MXFP4.

Terimlerden seçime

Yerel bir sohbet asistanı için kararların sırası oldukça sabittir. Instruct ya da reasoning varyantını alın, base modeli değil. Elinizdeki belleğe bakın ve ona sığan bir nicemleme seçin: Q4_K_M boyut ile kalite arasındaki olağan dengedir, Q8_0 neredeyse kayıpsızdır, Q3'ün altında ise yanıtlar belirgin biçimde bozulur. İhtiyacınız olan bağlama yer bırakın ve hızın kabul edilebilir olup olmadığına ancak bundan sonra karar verin. Bilinen modeller için bu hesabı donanım uyumu grafiği yapar.

Sık karıştırılan iki şey

Açık ağırlık, açık kaynakla aynı şey değildir. Dosyayı indirebilirsiniz, ancak lisans ticari kullanımı kısıtlayabilir ya da koşulları kabul etmenizi isteyebilir; eğitim verisi de genellikle kapalı kalır. Damıtılmış model de sık yanlış anlaşılır: büyük bir modelin yanıtlarıyla eğitilmiş daha küçük bir modeldir, büyük modelin kendisi değildir ve belirgin biçimde daha zayıftır.

Abliterated ve uncensored modeller hakkında

Bunlar, açık modellerden reddetme davranışı çıkarılmış topluluk değişiklikleridir. Terimi pek çok model adında geçtiği için listeliyoruz; bu tür modelleri önermiyoruz. Yalnızca kişisel ve araştırma amaçlı kullanım içindir ve nasıl kullanıldıklarının sorumluluğu kullanıcıya aittir.