Pixtral 12B 24.09 logo

Pixtral 12B 24.09オープンマルチモーダル 12B モデルが、128K コンテキストウィンドウでインターリーブ画像とテキストを処理する。

4.6 (5)
Daniel Nikulshynレビュー: Daniel Nikulshyn·更新 2026年7月

概要

Pixtral 12B 24.09はMistral AIから提供されるマルチモーダルモデルで、1つのシーケンス内で画像とテキストの両方を処理し、可変画像サイズとアスペクト比をサポートしています。 12億パラメータの言語デコーダを視覚エンコーダと組み合わせて使用すると、視覚質問応答、ドキュメント理解、チャート解釈、画像キャプションのタスクなどの機能が実現されます。 モデルは最大128Kトークンコンテキストを受け入れ、多数の画像を長文テキストと交互に提示したワードプромプを使うことを許可します。これにより、モデルは開発者が視覚言語アプリケーション、研究ワークフロー、および多モーダルエージェントを開発する際に、利用できるデータを幅広く拡大できます。 本モデルはオープンライセンスで提供され、ローカルにインストールできるほか、インファレンスプロバイダーを介したデプロイも可能です。

主な機能

  • 12Bパラメータのビジョン言語モデル
  • インターリーブ写真とテキスト入力
  • 128Kトークンコンテキスト長
  • ネイティブ変数画像サイズサポート
  • オープンウェイトリリース
  • OCR、VQA、キャプションツール用に適したもの

料金

モデル
Free
カテゴリー
カタトリアート
評価
4.6 / 5 (5)

ユースケース

マルチモダルトリーニング

Pixtral 12Bは自然画像とドキュメントを理解することができ、MMMU理屈基準で最先端パフォーマンスを達成しながら、より大きなモデルを上回る。

指示を遂行する

パストラル 12B は、多モダルおよびテキストのみのシナリオで特に優れており、IF-EvalおよびMT-Benchにおける最も近いオープンソースモデルに対する20%の相対的な改善を達成します。

マルチモダル質問への回答

pixtral 12B は強力な能力を持っているマルチモダル質問への回答を含め、ドキュメント質問回答、チャートおよびフィギュアの理解など。

メリット & デメリット

メリット

  • オープンウェイスト自宅デプロイ可能
  • 複数の写真 per プレームを処理
  • 大きな 128K コンテキストウィンドウ
  • 可変画像分解能とアスペクト比

デメリット

  • _GPUリソースを大量に必要とする
  • 閉鎖モデルよりも小さい
  • 限定工具比較API

バトル戦績

パンテオンで1バトルに出場。

0
1位
1
2位
0
3位

Last battle

レビュー

4.6

5件の評価の平均。

5
3
4
2
3
0
2
0
1
0

レビューを投稿するにはログインしてください。

SG

Sanjay Gupta

Jan 7, 2026

Does the job

Pretty happy overall. Open-weight release just works and large 128K context window. but no dealbreakers — I'd recommend it to a friend without hesitating.

Fatima Zahra

Fatima Zahra

Nov 26, 2025

Does the job

Pretty happy overall. Open-weight release just works and handles multiple images per prompt. but no dealbreakers — I'd recommend it to a friend without hesitating.

Naomi Suzuki

Naomi Suzuki

Oct 12, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is interleaved image and text inputs — handled better than most — and handles multiple images per prompt. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

TA

Tariq Aziz

Oct 7, 2025

Solid for our team

We rolled this out across the team last quarter and open weights for self-hosting. Open-weight release fits neatly into how we already work, and interleaved image and text inputs removed a step we used to do by hand. Smaller than frontier closed models, which is the main caveat, but it has held up under daily use.

Aaliyah Johnson

Aaliyah Johnson

Aug 30, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is 12B parameter vision-language model — handled better than most — and open weights for self-hosting. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

Q&A

How many images and how much text can I include in a single prompt?

Pixtral supports interleaved image and text inputs within a single 128 K token context, allowing any number of images (at their natural resolution) alongside long‑form text in one prompt.

Asked by Lorenzo Bianchi · Nov 20, 2025

Is Pixtral 12B still maintained, and are there newer alternatives?

Pixtral 12B is deprecated and no longer maintained; Mistral AI recommends using its newer, more powerful vision‑language models that supersede Pixtral for production use.

Asked by Carlos Mendoza · Nov 10, 2025

What hardware is needed to run Pixtral 12B effectively?

The model requires substantial GPU memory due to its 12 billion parameters and 400 M‑parameter vision encoder; typical deployments use high‑end GPUs (e.g., A100 40 GB or comparable) to handle the 128 K token context and multiple images.

Asked by Ivo Novotný · Nov 6, 2025

Can I self‑host Pixtral 12B, and under what license?

Yes, Pixtral 12B is released under the Apache 2.0 open‑source license, allowing you to download the weights and run the model locally on your own hardware.

Asked by Priya Nair · Oct 12, 2025

質問する

カタトリアートの代替