Ir al contenido
fedi.software

Tipos de modelos de IA: mapa y glosario

LLM o VLM, pesos abiertos o código abierto, base, instruct, reasoning, destilado, MoE, GGUF y cuantización — qué significan las palabras en los nombres de los modelos, con ejemplos en vivo.

Grupo Qué significa Ejemplos
LLMlarge language model · text model Tipo de modelo A model that reads and writes text: chat, writing, code, analysis. Everything else in this glossary is a flavour or a packaging of it.
VLMVL · vision-language · multimodal Tipo de modelo An LLM that also takes images (screenshots, photos, documents) as input. Locally it usually needs a second file — the vision projector (mmproj) — next to the GGUF weights.
Open weightsopen model · downloadable weights Apertura The weights can be downloaded and run on your own hardware, under any licence — some (Llama, Gemma) limit commercial use or require accepting terms first. Training data and code usually stay closed.
Open Source AIOSAID · OSI definition Apertura The stricter OSI definition (OSAID 1.0): weights plus the training code and enough information about the data to rebuild the model, all under open licences. Few models qualify; Apache-2.0 or MIT weights alone are open weights, not necessarily open source. —
Proprietary (API only)closed · API model Apertura No weights at all: the model runs only on the vendor's servers through an app or a paid API. It cannot be run locally. —
Basepretrained · -Base Variante The raw pretrained model: it continues text but does not follow instructions. A starting point for fine-tuning, not for chatting. —
Instruct-it · -Chat · -Instruct Variante The base model tuned to follow instructions and hold a dialogue. This is the version you want for a local chat; Google marks it -it, others -Instruct or -Chat.
ReasoningThinking · R1 · -Thinking Variante Trained to write out a chain of thought before the answer. Better at maths, code and logic, but spends many more tokens — and so more time — per reply.
Coder-Coder · code model Variante Further trained on source code: completion, refactoring, agentic coding in the IDE. General chat quality may be lower than the sibling instruct model.
Distill-Distill · R1-Distill-Qwen Variante A smaller model taught on the answers of a bigger one. DeepSeek-R1-Distill-Qwen-32B is a Qwen 32B that imitates R1 — not R1 itself, and far weaker than the 671B original. —
Abliterated / uncensoreduncensored · abliterated · heretic Sin filtros de seguridad Variante A community modification with the refusal behaviour removed from the weights. It answers anything, including harmful requests, often with lower quality and no safety guarantees. For personal and research use only; you are responsible for how you use it. —
Densedense model Arquitectura Every parameter works on every token. Speed is set by the full size: a 70B dense model reads all 70B weights from memory for each generated token.
MoEmixture of experts Arquitectura The layers are split into many experts and a router picks a few of them per token. All experts must sit in memory (VRAM + RAM), but each token touches only the active part — so a big MoE runs much faster than a dense model of the same size, and the experts can be offloaded to system RAM.
Active parameters (-A3B)-A3B · -A22B · active Arquitectura In MoE names the suffix gives the parameters used per token: Qwen3-30B-A3B has 30B in total and 3B active. Memory follows the total, speed follows the active number.
Safetensors (BF16)safetensors · BF16 · FP16 Formato de archivo The original release format on Hugging Face, usually 16 bits per weight: about 2 GB per billion parameters. Used by vLLM, Transformers and as the source for every quantisation. —
GGUF.gguf · llama.cpp Formato de archivo A single-file format of llama.cpp with the weights already quantised. Runs on CPU, GPU or both at once; LM Studio, Ollama and Jan use it. Big models come split into parts (-00001-of-00003). —
AWQ / GPTQ / EXL2-3AWQ · GPTQ · EXL2 · EXL3 Formato de archivo Quantised formats for GPU-only servers (vLLM, ExLlama, TGI). Fast when the whole model fits in VRAM; no offload to system RAM. —
MLXmlx-community Formato de archivo Apple's framework and weight format for M-series Macs, using the unified memory. Often a little faster on a Mac than GGUF of the same size. —
Q4_K_M and other K-quantsQ4_K_M · Q5_K_M · Q6_K · Q8_0 Cuantización GGUF quant names: the number is roughly the bits per weight, K marks the k-quant method, S/M/L the mix inside. Q4_K_M (≈4.8 bits) is the usual sweet spot; Q8_0 is almost lossless; below Q3 quality drops noticeably. —
IQ quants and Unsloth Dynamic (UD)IQ2_XXS · IQ3_K · UD-Q2_K_XL Cuantización Newer low-bit GGUF methods: IQ (importance-matrix) quants and Unsloth Dynamic keep the sensitive layers at higher precision, so 2–3-bit files of huge models stay usable.
FP8 / MXFP4 / NVFP4FP8 · MXFP4 · NVFP4 Cuantización Low-precision floating-point formats with hardware support in recent GPUs. gpt-oss ships natively in MXFP4 (≈4.25 bits per weight), which is why gpt-oss-120b fits in about 65 GB.

Los modelos marcados como «sin filtros de seguridad» (uncensored / abliterated) se incluyen solo para uso personal y de investigación: no tienen restricciones de seguridad y la responsabilidad de su uso recae en el usuario.

Actualizado · Fuentes: fedi.software (CC BY 4.0), Open Source Initiative (OSAID 1.0)

Insertar en tu sitio

Pega este código donde deba aparecer la infografía: se actualiza sola. Uso gratuito — conserva el enlace de atribución.

JSON Nuestros datos: CC BY 4.0 con enlace a la fuente; los datos de terceros mantienen la licencia de su fuente.

Cómo leer el nombre de un modelo

Los nombres de los modelos parecen claves, pero siguen casi siempre el mismo orden. Léelos por partes:

  1. familia y tamaño;
  2. en un MoE, los parámetros activos (-A3B);
  3. la variante (-Instruct o -it);
  4. por último, el formato de archivo y la cuantización de la descarga.

El glosario agrupa estos términos en seis bloques y enlaza cada uno con ejemplos reales en la comparación de modelos: “Tipo de modelo” (LLM, VLM), “Apertura” (pesos abiertos u open weights, Open Source AI según la definición 1.0 de la OSI, propietario), “Variante” (base, Instruct, reasoning, coder, destilado, abliterated / uncensored), “Arquitectura” (denso, MoE, parámetros activos), “Formato de archivo” (Safetensors, GGUF, AWQ / GPTQ / EXL, MLX) y “Cuantización” (K-quants como Q4_K_M, IQ y Unsloth Dynamic, FP8 / MXFP4).

Términos que se confunden

Pesos abiertos no equivale a código abierto. Puedes descargar los pesos, pero su licencia puede limitar el uso y los datos de entrenamiento suelen seguir cerrados. Un modelo destilado es un modelo más pequeño entrenado con las respuestas de otro mayor, no ese modelo grande. En cuantización, Q4_K_M es el equilibrio habitual entre tamaño y calidad, Q8_0 casi no pierde nada y por debajo de Q3 la calidad cae de forma notable.

Las variantes abliterated o uncensored tienen eliminado el comportamiento de rechazo. Aparecen de forma neutral y solo para uso personal y de investigación; la responsabilidad de su uso recae en el usuario y no son una recomendación.

Del glosario a un chat en tu equipo

Para montar un chat local, el camino es este: una variante Instruct (o reasoning), luego la memoria de la que dispones, una cuantización que quepa en ella, el contexto que necesitas y, al final, la velocidad que te parece aceptable. Los dos últimos pasos se comprueban en qué modelo abierto cabe en tu hardware. Las definiciones se basan en las fichas y configuraciones de Hugging Face y en la documentación de llama.cpp.