
Pixtral 12B 24.09Offenes multimodales 12B-Modell, das ineinander verschachtelte Bilder und Text mit einem Kontextfenster von 128 K verarbeitet.
Übersicht
Hauptfunktionen
- 12B-Parameter Vision‑Language‑Modell
- Eingänge aus ineinander verschachtelten Bild‑ und Textsequenzen
- 128 K Token Kontextlänge
- Native Unterstützung für variable Bildgrößen
- Open‑Weight‑Veröffentlichung
- Geeignet für OCR, VQA und Bildunterschrift
Preise
- Modell
- Free
- Kategorie
- Large Language Model
- Bewertung
- 4.6 / 5 (5)
Anwendungsfälle
Multimodales Denken
Pixtral 12B ist fähig, sowohl natürliche Bilder als auch Dokumente zu verstehen, erreicht state‑of‑the‑art Leistung auf dem MMMU Reasoning Benchmark und übertrifft größere Modelle.
Anweisungsbefolgung
Pixtral 12B überzeugt bei der Befolgung von Anweisungen, insbesondere in multimodalen und reinen Textszenarien, mit einer 20 %igen relativen Verbesserung bei text IF‑Eval und MT‑Bench gegenüber dem nächsten Open‑Source‑Modell.
Multimodale Fragebeantwortung
Pixtral 12B zeigt starke Fähigkeiten in multimodaler Fragebeantwortung, einschließlich Dokumentenfragebeantwortung sowie Diagramm- und Figurenverständnis.
Pro & Contra
Pro
- Offene Gewichte für Self‑Hosting
- Verarbeitet mehrere Bilder pro Prompt
- Großes 128 K Kontextfenster
- Flexible Bildauflösungen und Seitenverhältnisse
Contra
- Benötigt erhebliche GPU‑Ressourcen
- Kleiner als marktführende geschlossene Modelle
- Begrenzte Tooling im Vergleich zu proprietären APIs
Schlacht-Bilanz
Aus 1 Schlacht im Pantheon.
Last battle
Bewertungen
Durchschnitt aus 5 Bewertungen.
Melde dich an, um eine Bewertung abzugeben.
Does the job
Pretty happy overall. Open-weight release just works and large 128K context window. but no dealbreakers — I'd recommend it to a friend without hesitating.
Does the job
Pretty happy overall. Open-weight release just works and handles multiple images per prompt. but no dealbreakers — I'd recommend it to a friend without hesitating.
Years in this space
I've evaluated a lot of these over the years. What stands out here is interleaved image and text inputs — handled better than most — and handles multiple images per prompt. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.
Solid for our team
We rolled this out across the team last quarter and open weights for self-hosting. Open-weight release fits neatly into how we already work, and interleaved image and text inputs removed a step we used to do by hand. Smaller than frontier closed models, which is the main caveat, but it has held up under daily use.
Years in this space
I've evaluated a lot of these over the years. What stands out here is 12B parameter vision-language model — handled better than most — and open weights for self-hosting. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.
Fragen & Antworten
How many images and how much text can I include in a single prompt?
Pixtral supports interleaved image and text inputs within a single 128 K token context, allowing any number of images (at their natural resolution) alongside long‑form text in one prompt.
Asked by Lorenzo Bianchi · Nov 20, 2025
Is Pixtral 12B still maintained, and are there newer alternatives?
Pixtral 12B is deprecated and no longer maintained; Mistral AI recommends using its newer, more powerful vision‑language models that supersede Pixtral for production use.
Asked by Carlos Mendoza · Nov 10, 2025
What hardware is needed to run Pixtral 12B effectively?
The model requires substantial GPU memory due to its 12 billion parameters and 400 M‑parameter vision encoder; typical deployments use high‑end GPUs (e.g., A100 40 GB or comparable) to handle the 128 K token context and multiple images.
Asked by Ivo Novotný · Nov 6, 2025
Can I self‑host Pixtral 12B, and under what license?
Yes, Pixtral 12B is released under the Apache 2.0 open‑source license, allowing you to download the weights and run the model locally on your own hardware.
Asked by Priya Nair · Oct 12, 2025
Frage stellen
Alternativen zu Large Language Model

Hochleistungs-LLM-Gateway, das über 1000 Modelle hinter einer einzigen API vereint.

KI-Modell der nächsten Generation mit Fokus auf Schlussfolgerungen von DeepSeek

Open-Source-Mixture-of-Experts-Modell, das GPT-4o-ähnliche Denkfähigkeiten zu einem Bruchteil der Kosten bietet.

Conversational AI von xAI, entwickelt für logisches Denken, Forschung und Echtzeit-Antworten.

Meta's mehrsprachiges Open-Weight-LLM, abgestimmt auf effiziente, hochwertige Textgenerierung.

Auf der Basis von KI: MP3-zu-Text-Konverter zur Umwandlung von Audio in präzise lesbare Transkripte.

Ein quelloffenes großes Sprachmodell, das bei Denk-, Mathematik- und Codierungsaufgaben mit MIT-Lizenzierung für kostenlose Nutzung und Modifikation herausragt.

OpenAI's auf Logik ausgerichtetes Modell für komplexe, mehrstufige Problemlösungen.
Trending now

Genaue Hilfe bei Hausaufgaben mit ausführlichen Erklärungen

Intelligenter Dokument API zur Analyse, Trennung, OCR-Analyse und Strukturierung von komplexen PDFs, Präsentationen und Tabellenkalkulationen.

Gepflichtete Antworten mit Provision pro Klick.

Führender Plattform für agente Prozessautomatisierung, die sich selbst lernende AI-Agenten verwendet, um Arbeitsabläufe über verschiedene Branchen hinweg zu entzerren.
