أنواع نماذج الذكاء الاصطناعي: خريطة ومسرد
LLM أو VLM، أوزان مفتوحة أو مصدر مفتوح، base وinstruct وreasoning ومُقطَّر وMoE وGGUF والتكميم — ماذا تعني الكلمات في أسماء النماذج، مع أمثلة حية.
نوع النموذج
الانفتاح
البنية
صيغة الملف
التكميم
| المجموعة | ماذا يعني | أمثلة | |
|---|---|---|---|
| LLMlarge language model · text model | نوع النموذج | A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it. | |
| VLMVL · vision-language · multimodal | نوع النموذج | An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights. | |
| Open weightsopen model · downloadable weights | الانفتاح | The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed. | |
| Open Source AIOSAID · OSI definition | الانفتاح | The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. | — |
| Proprietary (API only)closed · API model | الانفتاح | No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. | — |
| Basepretrained · -Base | الإصدار | The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. | — |
| Instruct-it · -Chat · -Instruct | الإصدار | The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat. | |
| ReasoningThinking · R1 · -Thinking | الإصدار | Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply. | |
| Coder-Coder · code model | الإصدار | Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model. | |
| Distill-Distill · R1-Distill-Qwen | الإصدار | A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. | — |
| Abliterated / uncensoreduncensored · abliterated · heretic بدون مرشحات أمان | الإصدار | A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. | — |
| Densedense model | البنية | Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token. | |
| MoEmixture of experts | البنية | The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM. | |
| Active parameters (-A3B)-A3B · -A22B · active | البنية | In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number. | |
| Safetensors (BF16)safetensors · BF16 · FP16 | صيغة الملف | The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. | — |
| GGUF.gguf · llama.cpp | صيغة الملف | A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). | — |
| AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 | صيغة الملف | Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. | — |
| MLXmlx-community | صيغة الملف | Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. | — |
| Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 | التكميم | GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. | — |
| IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL | التكميم | Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable. | |
| FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 | التكميم | Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB. |
النماذج المشار إليها بعبارة “بدون مرشحات أمان” (uncensored / abliterated) مدرجة للاستخدام الشخصي والبحثي فقط: ليست لها قيود أمان، والمسؤولية عن استخدامها تقع على المستخدم.
تاريخ التحديث · المصادر: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)
تضمين في موقعك
الصق هذا الكود حيث يجب أن يظهر الإنفوغرافيك: يتحدّث تلقائيًا. الاستخدام مجاني — أبقِ رابط الإسناد.
كيف يُقرأ اسم النموذج؟
اسم مثل «Qwen3-30B-A3B-Instruct-GGUF Q4_K_M» يجمع عدة قرارات في سطر واحد. أولا العائلة والحجم الكلي، ثم في نموذج MoE المعاملات النشطة (-A3B)، ثم النسخة (-Instruct أو -it)، وأخيرا صيغة الملف ومستوى التكميم في التنزيل. المسرد في هذه الصفحة يقسّم هذه الكلمات إلى ست مجموعات، مع أمثلة حية بجانب كل مصطلح، والنقر على المثال يفتحه في مقارنة النماذج.
ما المجموعات الست؟
- نوع النموذج: LLM يعمل مع النص، وVLM يقبل الصور أيضا.
- الانفتاح: أوزان مفتوحة، أو تعريف Open Source AI الأكثر صرامة من OSI، أو نماذج مملوكة لا تُتاح إلا عبر تطبيق أو API.
- النسخة: base وInstruct وreasoning وcoder، والنماذج المقطّرة، ونسخ مجتمعية أزيل منها الرفض.
- البنية: النماذج الكثيفة مقابل MoE، ولاحقة المعاملات النشطة.
- صيغة الملف: Safetensors وGGUF، وصيغ خاصة بالمعالجات الرسومية مثل AWQ وGPTQ وEXL، وMLX لأجهزة Mac.
- التكميم: K-quants مثل Q4_K_M، وملفات IQ وUnsloth Dynamic، وFP8 وMXFP4.
بأي ترتيب أختار نموذجا لدردشة محلية؟
خذ نسخة Instruct أو reasoning لا نموذج base. ثم انظر إلى الذاكرة المتاحة لديك واختر تكميما يتسع لها: Q4_K_M هو التوازن المعتاد بين الحجم والجودة، وQ8_0 شبه خالٍ من الفقد، وتحت Q3 تتراجع الإجابات بوضوح. اترك مكانا للسياق الذي تحتاجه، وبعد ذلك فقط احكم على السرعة. هذا الحساب للنماذج المعروفة يجريه رسم ملاءمة العتاد.
هل الأوزان المفتوحة هي المصدر المفتوح؟
لا. يمكنك تنزيل الملف، لكن الترخيص قد يقيّد الاستخدام التجاري أو يطلب الموافقة على شروط، وبيانات التدريب تبقى مغلقة عادة. ويُساء فهم النموذج المقطّر (distilled) كذلك: هو نموذج أصغر دُرّب على إجابات نموذج أكبر، وليس النموذج الأكبر نفسه، وهو أضعف منه بوضوح.
ما معنى abliterated وuncensored؟
هي تعديلات مجتمعية على نماذج مفتوحة أزيل منها سلوك رفض الطلبات. نذكر المصطلح لأنه يتكرر في أسماء كثيرة، دون أن نوصي بهذه النماذج. وهي مخصصة للاستخدام الشخصي والبحثي فقط، والمسؤولية عن استخدامها تقع على المستخدم.