
Pixtral 12B 24.09여러 모드의 오픈 12B 모델은 128K 컨텍스트 윈도우가 있는 이미지를 처리하고 텍스트를 함께 사용합니다.
개요
주요 기능
- 12B 매개변수 비전-언어 모델
- 인터리브드 이미지와 텍스트 입력
- 128K 토큰 콘텍스트 길이
- ネ이티브 변수 사진 크기 지원
- 오픈 가중치 릴리스
- OCR, VQA 및 캡션을 위한 적합합니다.
가격
- 모델
- Free
- 카테고리
- 문이우지
- 평점
- 4.6 / 5 (5)
사용 사례
Multimodal Reasoning
Pixtral 12B는 자연 이미지를 포함하여 이해할 수 있으며 MMMU reasoning benchmark에서 최고 수준의 성능을 달성하고 더 큰 모델보다 능가합니다.
Instruction Following
Pixtral 12B는 특히 다중모드 및 텍스트のみ 대기 시나리오에서 지시를 따는데 뛰어나요, text IF-Eval 및 MT-Bench에서 가장 최근 오픈 소스 모델보다 20% 상대적으로 개선된 성능을 показ했습니다.
Multimodal Question Answering
Pixtral 12B는 문서 질문에 대한 답변과 차트 및 그림의 이해를 포함하여 다중 모드 질문에 강한 능력을 보입니다.
장단점
장점
- 자체 호스팅을 위한 오픈 가중치
- 여러 개의 이미지당 한 번의 프롬프트 처리
- _large 128K 콘텍스트 윈도우
- 가변적인 사진 해상도 및 측벽 비율
단점
- significnat GPU 리소스 요구
- 프론티어의閉鎖된 모델보다 더 작음
- proprietary API와 비교하여 제한된 도구
대결 기록
Pantheon에서 1회 대결.
Last battle
리뷰
5개 평가의 평균.
리뷰를 작성하려면 로그인하세요.
Does the job
Pretty happy overall. Open-weight release just works and large 128K context window. but no dealbreakers — I'd recommend it to a friend without hesitating.
Does the job
Pretty happy overall. Open-weight release just works and handles multiple images per prompt. but no dealbreakers — I'd recommend it to a friend without hesitating.
Years in this space
I've evaluated a lot of these over the years. What stands out here is interleaved image and text inputs — handled better than most — and handles multiple images per prompt. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.
Solid for our team
We rolled this out across the team last quarter and open weights for self-hosting. Open-weight release fits neatly into how we already work, and interleaved image and text inputs removed a step we used to do by hand. Smaller than frontier closed models, which is the main caveat, but it has held up under daily use.
Years in this space
I've evaluated a lot of these over the years. What stands out here is 12B parameter vision-language model — handled better than most — and open weights for self-hosting. Smaller than frontier closed models is my one real gripe. Worth the time if this is your use case.
Q&A
How many images and how much text can I include in a single prompt?
Pixtral supports interleaved image and text inputs within a single 128 K token context, allowing any number of images (at their natural resolution) alongside long‑form text in one prompt.
Asked by Lorenzo Bianchi · Nov 20, 2025
Pixtral 12B는 여전히 유지 관리되고 있나요? 더 최신의 대안이 있나요?
Pixtral 12B는 더 이상 유지 관리되지 않으며, Mistral AI는 Pixtral을 대체하는 보다 강력한 비전‑언어 모델을 생산 환경에서 사용하도록 권장합니다.
Asked by Carlos Mendoza · Nov 10, 2025
Pixtral 12B를 효과적으로 실행하려면 어떤 하드웨어가 필요합니까?
이 모델은 120억 개 매개변수와 400 M‑parameter 비전 인코더 때문에 상당한 GPU 메모리를 요구합니다. 일반적으로 A100 40 GB 또는 이와 상응하는 고급 GPU를 사용하여 128 K 토큰 컨텍스트와 다중 이미지를 처리합니다.
Asked by Ivo Novotný · Nov 6, 2025
Pixtral 12B를 자체 호스팅할 수 있나요? 라이선스는 어떻게 되나요?
네, Pixtral 12B는 Apache 2.0 오픈소스 라이선스 하에 배포되며, 가중치를 다운로드 받아 자체 하드웨어에서 모델을 실행할 수 있습니다.
Asked by Priya Nair · Oct 12, 2025
질문하기
문이우지 대안

1000+

다음세대 사고를 중심으로 한 인공지능 모델: DeepSeek

오픈 소스 mixture-of-experts 모델로 GPT-4 단계의 주관호적 사고 능력 1/4의 비용만 내면 됩니다.

xAI가 개발 한 논리적 reasoning, 연구 및 실시간 답변을 위한 대화형 AI.

효율적이고 고품질의 텍스트 생성을 위한 멀티 언어 개방 가중치의 LLM

음성 파일을 깨끗하고 읽을 수 있는 전사본으로 바꾸는 AI-powered MP3 변환기

논리 추론, 수학, 프로그래밍 업무에 있어서 전문성을 가지는 오픈 소스 인공지능 모델이며 MIT 라이선스 아래 자유롭게 사용 및 수정이 가능합니다.

OpenAI가 개발한 복잡한 다단계 문제해결을 위한 의도력 있는 모델.
Trending now

정확한 과제 도움말과 전체 설명

복잡한 PDF, 슬라이드 및 스프레드시트에서 구조화된 데이터를 추출할 수 있는 문서 지능 API입니다. 이들은 텍스트, 이미지, 탭 및 레이아웃 정보까지 모든 요소를 파싱 및 추출합니다.

광고 주도 답변, 클릭당 지불

각 산업에서 워크플로를 최적화하기 위한 리더적인 에이전트 프로세스 자동화 플랫폼으로, 자체 학습 에이전트를 활용하여 다중 산업의 워크플로우를 개선한다.
