Phi-4 Multimodal Instruct
MicrosoftPhiオープンウエイトMIT · 商用利用可
説明
Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.
リリース日
2025-02-26
パラメータ
5.6B
コンテキスト長
—
モダリティ
image, text
能力レーダー
18
general
13
coding
32
reasoning
19
science
26
agents
85
multimodal
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| コーディングランキング | 467 | 14.0 | AA |
| 総合ランキング | 500 | 20.0 | AA |
| マルチモーダルランキング | 74 | 29.0 | LS |
| 科学 | 508 | 17.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))3d
BLINK
61.3%自己申告
General
MMMU
55.1%自己申告
MMMU-Pro
38.5%自己申告
Image To Text
DocVQADocVQA (2020)
93.2%自己申告
OCRBench
84.4%自己申告
TextVQA
75.6%自己申告
Math
MathVista
62.4%自己申告
InterGPS
48.6%自己申告
Multimodal
ScienceQA Visual
97.5%自己申告
MMBench
86.7%自己申告
POPE
85.6%自己申告
AI2D
82.3%自己申告
ChartQAMasry et al. (2022)
81.4%自己申告
InfoVQA
72.7%自己申告
Video-MME
55.0%自己申告
AA評価指数
(Artificial Analysis)Intelligence Index(Artificial Analysis)4.2
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))0.7
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))0.5
Gpqa(NYU + Cohere + Anthropic (2023))0.3
Livecodebench(UC Berkeley + MIT + Cornell (2024))0.1
Scicode(UIUC + Argonne National Lab (2024))0.1
Aime(MAA (Mathematical Association of America))0.1
Hle(Center for AI Safety + Scale AI (2025))0.1
LLM Statsカテゴリスコア
(LLM Stats (zeroeval))Safety90
Image To Text80
Multimodal70
Reasoning70
General70
Vision70
Math60
Spatial Reasoning60
Healthcare60
3d60
価格設定
入力価格無料
出力価格無料
混合価格(3:1)無料
速度
トークン/秒17.7
初トークン遅延0.37s
初回答遅延0.37s
プロバイダー価格ランキング
プロバイダー価格ランキング
2 プロバイダー
最安: Azure Cognitive Services最高: Azure
プロバイダー入力出力
1Azure Cognitive Services最安
$0.08
$0.32
2Azure
$0.08
$0.32
このモデルの異なるAPIプロバイダー間の価格を比較。