Pixtral 12B 24.09 logo

Pixtral 12B 24.09Çıkarılmış modda 12B model, 128K bağlam penceresi ile karışık görüntü ve metin ile işlenir.

4.6 (5)
Daniel Nikulshynİnceleyen Daniel Nikulshyn·Güncellendi Temmuz 2026

Genel Bakış

Pixtral 12B 24.09, Mistral AI'den bir modası geçen görsel-modali bir modeldir. Bir tek dizi içindeki resim ve metni işler ve varyasyonlu görüntü boyutları ve aspect oranları da destekler. 12 milyar parametreye sahip bir dil kodlama cihazını görme kodlama cihazıyla birleştiren Pixtral 12B 24.09, görsel soruları yanıtlamak, belge anlama, grafik yorumlama, görüntü başlıkları oluşturma gibi görevleri destekler. Model, 128K token altyapısını destekleyerek, çoklu görüntü ile uzun metne entegre bir komut istemi oluşturmak için birçok imgeyi bir komut istemi içinde birleştirmeyi sağlar. Açık lisans altında yayınlanan bu model, yerel olarak veya ön-yargı sağlayıcıları aracılığıyla kullanılabilecek şekilde dağıtımdan geçirilebilir, bu nedenle görüntü-dil uygulamaları oluşturmak için geliştiricilere, araştırmaya yönelik akışlar ve çok kanallı ajanlar inşa için elverişlidir.

Temel özellikler

  • 12B parametrenli bir görü-görü-dil modelidir.
  • İlişkili resim ve metin girişleri desteği.
  • 128K token bağlama uzunluğu.
  • Oluşkan resim boyutu desteği.
  • Açık ağırlık sürümü.
  • OCR, VQA ve etiketleme gibi uygulamalar için uygundur.

Fiyatlar

Model
Free
Kategori
LLM
Puan
4.6 / 5 (5)

Kullanım senaryoları

Çokmodal Mantıksal Çözümleme

Pixtral 12B, doğal resim ve belgeleri anlama kabiliyeti ile MMMU mantıksal çözümleme testi'nde mükemmele ulaşan ve daha büyük modellerden daha iyi sonuçlar isteyen bir modeldir.

Temsil İşleme

Pixtral 12B, özellikle çokmodal ve metin sadece senaryolarda, text IF-Eval ve MT-Bench'te 20'lük relative iyileşme ile daha yakın açık kaynak modelinden iyi performans gösteriyor.

Çokmodal Soru-Cevap

Pixtral 12B, doküman soru-cevap ve grafik ve şekil anlama konusunda güçlü abilities'e sahiptir.

Artılar ve eksiler

Artılar

  • Açık ağırlıklar için kendi sunucusunda barınma imkanı
  • Prompt'ta birden fazla görüntü işleme
  • Büyük 128K bağlama penceresi
  • Esnek resim çözünürlükleri ve açıları desteği

Eksiler

  • Önemli GPU kaynağı gerektirir
  • Cephe sınırı açık modellerinin daha küçüğüdür
  • Proprietary API'lerden daha az araç desteği

Savaş rekoru

Pantheon’da 1 savaş.

0
1.
1
2.
0
3.

Last battle

İncelemeler

4.6

5 puandan ortalama.

5
3
4
2
3
0
2
0
1
0

İnceleme bırakmak için giriş yap.

SG

Sanjay Gupta

Jan 7, 2026

Does the job

Pretty happy overall. Open-weight release just works and large 128K context window. but no dealbreakers — I'd recommend it to a friend without hesitating.

Fatima Zahra

Fatima Zahra

Nov 26, 2025

Does the job

Pretty happy overall. Open-weight release just works and handles multiple images per prompt. but no dealbreakers — I'd recommend it to a friend without hesitating.

Naomi Suzuki

Naomi Suzuki

Oct 12, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is interleaved image and text inputs — handled better than most — and handles multiple images per prompt. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

TA

Tariq Aziz

Oct 7, 2025

Solid for our team

We rolled this out across the team last quarter and open weights for self-hosting. Open-weight release fits neatly into how we already work, and interleaved image and text inputs removed a step we used to do by hand. Smaller than frontier closed models, which is the main caveat, but it has held up under daily use.

Aaliyah Johnson

Aaliyah Johnson

Aug 30, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is 12B parameter vision-language model — handled better than most — and open weights for self-hosting. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

Sorular

How many images and how much text can I include in a single prompt?

Pixtral supports interleaved image and text inputs within a single 128 K token context, allowing any number of images (at their natural resolution) alongside long‑form text in one prompt.

Asked by Lorenzo Bianchi · Nov 20, 2025

Is Pixtral 12B still maintained, and are there newer alternatives?

Pixtral 12B is deprecated and no longer maintained; Mistral AI recommends using its newer, more powerful vision‑language models that supersede Pixtral for production use.

Asked by Carlos Mendoza · Nov 10, 2025

What hardware is needed to run Pixtral 12B effectively?

The model requires substantial GPU memory due to its 12 billion parameters and 400 M‑parameter vision encoder; typical deployments use high‑end GPUs (e.g., A100 40 GB or comparable) to handle the 128 K token context and multiple images.

Asked by Ivo Novotný · Nov 6, 2025

Can I self‑host Pixtral 12B, and under what license?

Yes, Pixtral 12B is released under the Apache 2.0 open‑source license, allowing you to download the weights and run the model locally on your own hardware.

Asked by Priya Nair · Oct 12, 2025

Soru sor

LLM alternatifleri