Phi-4 Multimodal Instruct
MicrosoftPhi오픈 웨이트MIT · 상업적 사용 가능
설명
Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.
출시일
2025-02-26
파라미터
5.6B
컨텍스트 길이
—
모달리티
image, text
능력 레이더
18
general
13
coding
32
reasoning
19
science
26
agents
85
multimodal
랭킹
벤치마크 점수 (LLM Stats)
(LLM Stats (zeroeval))3d
BLINK
61.3%자체 보고
General
MMMU
55.1%자체 보고
MMMU-Pro
38.5%자체 보고
Image To Text
DocVQADocVQA (2020)
93.2%자체 보고
OCRBench
84.4%자체 보고
TextVQA
75.6%자체 보고
Math
MathVista
62.4%자체 보고
InterGPS
48.6%자체 보고
Multimodal
ScienceQA Visual
97.5%자체 보고
MMBench
86.7%자체 보고
POPE
85.6%자체 보고
AI2D
82.3%자체 보고
ChartQAMasry et al. (2022)
81.4%자체 보고
InfoVQA
72.7%자체 보고
Video-MME
55.0%자체 보고
AA 평가 지수
(Artificial Analysis)Intelligence Index(Artificial Analysis)4.2
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))0.7
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))0.5
Gpqa(NYU + Cohere + Anthropic (2023))0.3
Livecodebench(UC Berkeley + MIT + Cornell (2024))0.1
Scicode(UIUC + Argonne National Lab (2024))0.1
Aime(MAA (Mathematical Association of America))0.1
Hle(Center for AI Safety + Scale AI (2025))0.1
LLM Stats 카테고리 점수
(LLM Stats (zeroeval))Image To Text80
Multimodal70
Reasoning70
General70
Vision70
Math60
Spatial Reasoning60
Healthcare60
3d60
가격
입력 가격무료
출력 가격무료
혼합 가격 (3:1)무료
속도
토큰/초17.8
첫 토큰 지연0.35s
첫 응답 지연0.35s
공급자 가격 순위
공급자 가격 순위
2개 공급자
최저가: Azure Cognitive Services최고가: Azure
공급자입력출력
1Azure Cognitive Services최저가
$0.08
$0.32
2Azure
$0.08
$0.32
이 모델의 다양한 API 공급자 간 가격 비교.