Pixtral 12B 24.09 logo

Pixtral 12B 24.09여러 모드의 오픈 12B 모델은 128K 컨텍스트 윈도우가 있는 이미지를 처리하고 텍스트를 함께 사용합니다.

4.6 (5)
Daniel Nikulshyn리뷰어 Daniel Nikulshyn·업데이트됨 2026년 7월

개요

Pixtral 12B 24..09은 Mistral AI에서 제공하는 멀티 모달 모델로 이미지와 텍스트를 포함한 단일 시퀀스에서 처리할 수 있습니다. 다양한 이미지 크기และสัดส조를 지원합니다. 그것은 12억 매개변수의 언어 디코더와 비전 인코더의 조합을 사용하여, 시각적인 질문 답변, 문서 이해, 막대 그래프 해석, 이미지 캡션화 등과 같은 작업이 가능합니다. 이 모델은 최대 128K 토큰 컨텍스트를 받을 수 있다. 여러 이미지를 긴 텍스트와 교차하여 한 번에 하나의 명령으로 전달할 수 있게 한다. opensource 라이센스 하에 출시되어, 로컬에서 또는 인퍼런스 프로바이더를 통해 배포할 수 있어서, 개발자들이 시각-언어 어플리케이션, 연구 워크플로우 및 멀티모달 에이전트를 빌드해야 하는 데 적합하다.

주요 기능

  • 12B 매개변수 비전-언어 모델
  • 인터리브드 이미지와 텍스트 입력
  • 128K 토큰 콘텍스트 길이
  • ネ이티브 변수 사진 크기 지원
  • 오픈 가중치 릴리스
  • OCR, VQA 및 캡션을 위한 적합합니다.

가격

모델
Free
카테고리
문이우지
평점
4.6 / 5 (5)

사용 사례

Multimodal Reasoning

Pixtral 12B는 자연 이미지를 포함하여 이해할 수 있으며 MMMU reasoning benchmark에서 최고 수준의 성능을 달성하고 더 큰 모델보다 능가합니다.

Instruction Following

Pixtral 12B는 특히 다중모드 및 텍스트のみ 대기 시나리오에서 지시를 따는데 뛰어나요, text IF-Eval 및 MT-Bench에서 가장 최근 오픈 소스 모델보다 20% 상대적으로 개선된 성능을 показ했습니다.

Multimodal Question Answering

Pixtral 12B는 문서 질문에 대한 답변과 차트 및 그림의 이해를 포함하여 다중 모드 질문에 강한 능력을 보입니다.

장단점

장점

  • 자체 호스팅을 위한 오픈 가중치
  • 여러 개의 이미지당 한 번의 프롬프트 처리
  • _large 128K 콘텍스트 윈도우
  • 가변적인 사진 해상도 및 측벽 비율

단점

  • significnat GPU 리소스 요구
  • 프론티어의閉鎖된 모델보다 더 작음
  • proprietary API와 비교하여 제한된 도구

대결 기록

Pantheon에서 1회 대결.

0
1위
1
2위
0
3위

Last battle

리뷰

4.6

5개 평가의 평균.

5
3
4
2
3
0
2
0
1
0

리뷰를 작성하려면 로그인하세요.

SG

Sanjay Gupta

Jan 7, 2026

Does the job

Pretty happy overall. Open-weight release just works and large 128K context window. but no dealbreakers — I'd recommend it to a friend without hesitating.

Fatima Zahra

Fatima Zahra

Nov 26, 2025

Does the job

Pretty happy overall. Open-weight release just works and handles multiple images per prompt. but no dealbreakers — I'd recommend it to a friend without hesitating.

Naomi Suzuki

Naomi Suzuki

Oct 12, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is interleaved image and text inputs — handled better than most — and handles multiple images per prompt. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

TA

Tariq Aziz

Oct 7, 2025

Solid for our team

We rolled this out across the team last quarter and open weights for self-hosting. Open-weight release fits neatly into how we already work, and interleaved image and text inputs removed a step we used to do by hand. Smaller than frontier closed models, which is the main caveat, but it has held up under daily use.

Aaliyah Johnson

Aaliyah Johnson

Aug 30, 2025

Years in this space

I've evaluated a lot of these over the years. What stands out here is 12B parameter vision-language model — handled better than most — and open weights for self-hosting. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.

Q&A

How many images and how much text can I include in a single prompt?

Pixtral supports interleaved image and text inputs within a single 128 K token context, allowing any number of images (at their natural resolution) alongside long‑form text in one prompt.

Asked by Lorenzo Bianchi · Nov 20, 2025

Pixtral 12B는 여전히 유지 관리되고 있나요? 더 최신의 대안이 있나요?

Pixtral 12B는 더 이상 유지 관리되지 않으며, Mistral AI는 Pixtral을 대체하는 보다 강력한 비전‑언어 모델을 생산 환경에서 사용하도록 권장합니다.

Asked by Carlos Mendoza · Nov 10, 2025

Pixtral 12B를 효과적으로 실행하려면 어떤 하드웨어가 필요합니까?

이 모델은 120억 개 매개변수와 400 M‑parameter 비전 인코더 때문에 상당한 GPU 메모리를 요구합니다. 일반적으로 A100 40 GB 또는 이와 상응하는 고급 GPU를 사용하여 128 K 토큰 컨텍스트와 다중 이미지를 처리합니다.

Asked by Ivo Novotný · Nov 6, 2025

Pixtral 12B를 자체 호스팅할 수 있나요? 라이선스는 어떻게 되나요?

네, Pixtral 12B는 Apache 2.0 오픈소스 라이선스 하에 배포되며, 가중치를 다운로드 받아 자체 하드웨어에서 모델을 실행할 수 있습니다.

Asked by Priya Nair · Oct 12, 2025

질문하기

문이우지 대안