Pixtral 12B 24.09 logo

Pixtral 12B 24.09Modelo multimodal de 12B aberto que lida com imagens e texto intercalados com uma janela de contexto de 128K.

4.6 (5)
Daniel NikulshynAvaliado por Daniel Nikulshyn·Atualizado julho de 2026

Visão geral

Pixtral 12B 24.09 é um modelo multimodal da Mistral AI que processa tanto imagens quanto texto dentro de uma única sequência, suportando tamanhos de imagem variáveis e proporções. Ele usa um decodificador de linguagem com 12 bilhões de parâmetros emparelhado com um codificador de visão, possibilitando tarefas como responder a perguntas visuais, compreender documentos, interpretar gráficos e legendas de imagens. O modelo aceita até um contexto de 128K tokens, permitindo que várias imagens sejam intercaladas com texto de longo formato em uma única solicitação. Lançado sob uma licença aberta, pode ser implantado localmente ou por meio de provedores de inferência, tornando-o adequado para desenvolvedores que constroem aplicações de visão-linguagem, fluxos de trabalho de pesquisa e agentes multimodais.

Funcionalidades principais

  • Modelo de visão-linguagem com 12B parâmetros
  • Entradas de imagem e texto intercaladas
  • Comprimento de contexto de 128K tokens
  • Suporte nativo a tamanho de imagem variável
  • Lançamento com pesos abertos
  • Adequado para OCR, VQA e legendas

Preços

Modelo
Free
Avaliação
4.6 / 5 (5)

Casos de uso

Raciocínio Multimodal

O Pixtral 12B é capaz de compreender tanto imagens naturais quanto documentos, alcançando desempenho de última geração no benchmark de raciocínio MMMU e superando modelos maiores.

Seguir Instruções

O Pixtral 12B se destaca em seguir instruções, particularmente em cenários multimodais e apenas de texto, com uma melhoria relativa de 20% no texto IF-Eval e MT-Bench em relação ao modelo de código aberto mais próximo.

Perguntas e Respostas Multimodais

O Pixtral 12B mostra fortes habilidades em perguntas e respostas multimodais, incluindo perguntas e respostas de documentos e compreensão de gráficos e figuras.

Prós e contras

Prós

  • Pesos abertos para auto-hospedagem
  • Lida com várias imagens por solicitação
  • Janela de contexto grande de 128K
  • Resoluções de imagem e proporções flexíveis

Contras

  • Requer recursos significativos de GPU
  • Menor que modelos fechados de ponta
  • Ferramentas limitadas em comparação com APIs proprietárias

Histórico de batalhas

Em 1 batalha no Panteão.

0
1.º
1
2.º
0
3.º

Last battle

Avaliações

4.6

Média de 5 avaliações.

5
3
4
2
3
0
2
0
1
0

Entra para deixar uma avaliação.

SG

Sanjay Gupta

Jan 7, 2026

Does the job

Pretty happy overall. Open-weight release just works and large 128K context window. but no dealbreakers — I'd recommend it to a friend without hesitating.

Fatima Zahra

Fatima Zahra

Nov 26, 2025

Does the job

Pretty happy overall. Open-weight release just works and handles multiple images per prompt. but no dealbreakers — I'd recommend it to a friend without hesitating.

Naomi Suzuki

Naomi Suzuki

Oct 12, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is interleaved image and text inputs — handled better than most — and handles multiple images per prompt. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

TA

Tariq Aziz

Oct 7, 2025

Solid for our team

We rolled this out across the team last quarter and open weights for self-hosting. Open-weight release fits neatly into how we already work, and interleaved image and text inputs removed a step we used to do by hand. Smaller than frontier closed models, which is the main caveat, but it has held up under daily use.

Aaliyah Johnson

Aaliyah Johnson

Aug 30, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is 12B parameter vision-language model — handled better than most — and open weights for self-hosting. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

Perguntas e respostas

How many images and how much text can I include in a single prompt?

Pixtral supports interleaved image and text inputs within a single 128 K token context, allowing any number of images (at their natural resolution) alongside long‑form text in one prompt.

Asked by Lorenzo Bianchi · Nov 20, 2025

Is Pixtral 12B still maintained, and are there newer alternatives?

Pixtral 12B is deprecated and no longer maintained; Mistral AI recommends using its newer, more powerful vision‑language models that supersede Pixtral for production use.

Asked by Carlos Mendoza · Nov 10, 2025

What hardware is needed to run Pixtral 12B effectively?

The model requires substantial GPU memory due to its 12 billion parameters and 400 M‑parameter vision encoder; typical deployments use high‑end GPUs (e.g., A100 40 GB or comparable) to handle the 128 K token context and multiple images.

Asked by Ivo Novotný · Nov 6, 2025

Can I self‑host Pixtral 12B, and under what license?

Yes, Pixtral 12B is released under the Apache 2.0 open‑source license, allowing you to download the weights and run the model locally on your own hardware.

Asked by Priya Nair · Oct 12, 2025

Faz uma pergunta

Alternativas a ML de Nível Superior