
Pixtral 12B 24.09Open multimodal 12B model handling interleaved images and text with a 128K context window.
Overview
Key features
- 12B parameter vision-language model
- Interleaved image and text inputs
- 128K token context length
- Native variable image size support
- Open-weight release
- Suitable for OCR, VQA, and captioning
Pricing
- Model
- Free
- Category
- LLM
- Rating
- 4.6 / 5 (5)
Use cases
Multimodal Reasoning
Pixtral 12B is capable of understanding both natural images and documents, achieving state-of-the-art performance on the MMMU reasoning benchmark and surpassing larger models.
Instruction Following
Pixtral 12B excels in instruction following, particularly in multimodal and text-only scenarios, with a 20% relative improvement in text IF-Eval and MT-Bench over the nearest open-source model.
Multimodal Question Answering
Pixtral 12B shows strong abilities in multimodal question answering, including document question answering and chart and figure understanding.
Pros & Cons
Pros
- Open weights for self-hosting
- Handles multiple images per prompt
- Large 128K context window
- Flexible image resolutions and aspect ratios
Cons
- Requires significant GPU resources
- Smaller than frontier closed models
- Limited tooling compared to proprietary APIs
Battle record
Across 1 battle in the Pantheon.
Last battle
Reviews
Average from 5 ratings.
Sign in to leave a review.
Does the job
Pretty happy overall. Open-weight release just works and large 128K context window. but no dealbreakers — I'd recommend it to a friend without hesitating.
Does the job
Pretty happy overall. Open-weight release just works and handles multiple images per prompt. but no dealbreakers — I'd recommend it to a friend without hesitating.
Years in this space
I've evaluated a lot of these over the years. What stands out here is interleaved image and text inputs — handled better than most — and handles multiple images per prompt. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.
Solid for our team
We rolled this out across the team last quarter and open weights for self-hosting. Open-weight release fits neatly into how we already work, and interleaved image and text inputs removed a step we used to do by hand. Smaller than frontier closed models, which is the main caveat, but it has held up under daily use.
Years in this space
I've evaluated a lot of these over the years. What stands out here is 12B parameter vision-language model — handled better than most — and open weights for self-hosting. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.
Q&A
How many images and how much text can I include in a single prompt?
Pixtral supports interleaved image and text inputs within a single 128 K token context, allowing any number of images (at their natural resolution) alongside long‑form text in one prompt.
Asked by Lorenzo Bianchi · Nov 20, 2025
Is Pixtral 12B still maintained, and are there newer alternatives?
Pixtral 12B is deprecated and no longer maintained; Mistral AI recommends using its newer, more powerful vision‑language models that supersede Pixtral for production use.
Asked by Carlos Mendoza · Nov 10, 2025
What hardware is needed to run Pixtral 12B effectively?
The model requires substantial GPU memory due to its 12 billion parameters and 400 M‑parameter vision encoder; typical deployments use high‑end GPUs (e.g., A100 40 GB or comparable) to handle the 128 K token context and multiple images.
Asked by Ivo Novotný · Nov 6, 2025
Can I self‑host Pixtral 12B, and under what license?
Yes, Pixtral 12B is released under the Apache 2.0 open‑source license, allowing you to download the weights and run the model locally on your own hardware.
Asked by Priya Nair · Oct 12, 2025
Ask a question
LLM alternatives

High-performance LLM gateway unifying 1000+ models behind a single API.

Next-generation reasoning-focused AI model from DeepSeek

Open-source mixture-of-experts model offering GPT-4o-level reasoning at a fraction of the cost.

Conversational AI from xAI built for reasoning, research, and real-time answers.

Meta's multilingual open-weight LLM tuned for efficient, high-quality text generation.

AI-powered MP3 to text converter for turning audio into clean, readable transcripts.

An open-source large language model excelling in reasoning, math, and coding tasks with MIT licensing for free use and modification.

OpenAI's reasoning-focused model built for complex, multi-step problem solving.
Trending now

Accurate Homework Help with Full Explanations

Document intelligence API that parses, splits, OCRs, and extracts structured data from complex PDFs, slides, and spreadsheets.

Sponsored answers, paid per click.

A leading platform for agentic process automation, utilizing self-learning AI agents to streamline workflows across various industries.
