Pixtral 12B 24.09 logo

Pixtral 12B 24.09Modelo multimodal de 12G que maneja imágenes e texto intercalados con una ventana de contexto de 128K.

4.6 (5)
Daniel NikulshynReseñado por Daniel Nikulshyn·Actualizado julio de 2026

Resumen

Pixtral 12B 24.09 es un modelo multimodal de Mistral AI que procesa tanto imágenes como texto dentro de una sola secuencia, admitiendo tamaños de imagen y relaciones de aspecto variables. Utiliza un decodificador de lenguaje de 12 mil millones de parámetros emparejado con un codificador de visión, lo que permite tareas como responder preguntas visuales, comprender documentos, interpretar gráficos y subtitular imágenes. El modelo acepta un contexto de hasta 128K tokens, lo que permite intercalar varias imágenes con texto de formato largo en una sola solicitud. Publicado bajo una licencia abierta, se puede implementar de forma local o a través de proveedores de inferencia, lo que lo hace adecuado para desarrolladores que crean aplicaciones de lenguaje visual, flujos de trabajo de investigación y agentes multimodales.

Funciones clave

  • Modelo de visión-lenguaje de 12 mil millones de parámetros
  • Entradas de imágenes y texto intercalados
  • Longitud de contexto de 128K tokens
  • Soporte nativo de tamaños de imagen variables
  • Publicación de pesos abierta
  • Procedente para UO, PVA y etiquetado

Precio

Modelo
Free
Valoración
4.6 / 5 (5)

Casos de uso

Razonamiento Multimodal

Pixtral 12B es capaz de comprender tanto imágenes naturales como documentos, logrando el rendimiento de estado del arte en el benchmark de razonamiento MMMU y superando a modelos más grandes.

Seguimiento de Instrucciones

Pixtral 12B excellece en el seguimiento de instrucciones, en particular en escenarios multimodal y con texto solo, con un aumento relativo del 20% en el seguimiento de instrucciones IF-Eval y MT-Bench sobre el modelo abierto más cercano.

Preguntas Responda Multimodal

Pixtral 12B muestra fuertes habilidades para preguntas y respuestas multimodales, incluyendo la comprensión de preguntas-documento y comprensión de gráficos e imágenes.

Pros y contras

Ventajas

  • Peso abierto para hospedaje propio
  • Gestión de múltiples imágenes por síntesis
  • Gran ventana de contexto de 128K
  • Resoluciones y aspectos de imagen flexibles

Contras

  • Requiere recursos de GPU muy significativos
  • Pequeño en comparación con modelos cerrados de vanguardia
  • Limitaciones en herramientas frente a APIs propietarias

Historial de batallas

En 1 batalla del Panteón.

0
1.º
1
2.º
0
3.º

Last battle

Reseñas

4.6

Promedio de 5 valoraciones.

5
3
4
2
3
0
2
0
1
0

Inicia sesión para dejar una reseña.

SG

Sanjay Gupta

Jan 7, 2026

Does the job

Pretty happy overall. Open-weight release just works and large 128K context window. but no dealbreakers — I'd recommend it to a friend without hesitating.

Fatima Zahra

Fatima Zahra

Nov 26, 2025

Does the job

Pretty happy overall. Open-weight release just works and handles multiple images per prompt. but no dealbreakers — I'd recommend it to a friend without hesitating.

Naomi Suzuki

Naomi Suzuki

Oct 12, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is interleaved image and text inputs — handled better than most — and handles multiple images per prompt. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

TA

Tariq Aziz

Oct 7, 2025

Solid for our team

We rolled this out across the team last quarter and open weights for self-hosting. Open-weight release fits neatly into how we already work, and interleaved image and text inputs removed a step we used to do by hand. Smaller than frontier closed models, which is the main caveat, but it has held up under daily use.

Aaliyah Johnson

Aaliyah Johnson

Aug 30, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is 12B parameter vision-language model — handled better than most — and open weights for self-hosting. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

Preguntas y respuestas

How many images and how much text can I include in a single prompt?

Pixtral supports interleaved image and text inputs within a single 128 K token context, allowing any number of images (at their natural resolution) alongside long‑form text in one prompt.

Asked by Lorenzo Bianchi · Nov 20, 2025

Is Pixtral 12B still maintained, and are there newer alternatives?

Pixtral 12B is deprecated and no longer maintained; Mistral AI recommends using its newer, more powerful vision‑language models that supersede Pixtral for production use.

Asked by Carlos Mendoza · Nov 10, 2025

What hardware is needed to run Pixtral 12B effectively?

The model requires substantial GPU memory due to its 12 billion parameters and 400 M‑parameter vision encoder; typical deployments use high‑end GPUs (e.g., A100 40 GB or comparable) to handle the 128 K token context and multiple images.

Asked by Ivo Novotný · Nov 6, 2025

Can I self‑host Pixtral 12B, and under what license?

Yes, Pixtral 12B is released under the Apache 2.0 open‑source license, allowing you to download the weights and run the model locally on your own hardware.

Asked by Priya Nair · Oct 12, 2025

Hacer una pregunta

Alternativas a Modelos de lenguaje por largo contexto (LLM)