Phi-4 Multimodal Instruct
MicrosoftPhiオープンウエイトMIT · 商用利用可
説明
Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.
リリース日
2025-02-26
パラメータ
5.6B
コンテキスト長
—
モダリティ
image, text
能力レーダー
18
general
13
coding
32
reasoning
23
science
26
agents
85
multimodal
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| コーディングランキング | 561 | 14.0 | AA |
| 総合ランキング | 582 | 20.0 | AA |
| マルチモーダルランキング | 141 | 23.0 | LS |
| 科学 | 605 | 16.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Math
MathVista
62.4%自己申告
InterGPS
48.6%自己申告
Multimodal
MMMU
55.1%自己申告
Video-MME
55.0%自己申告
Reasoning
ChartQAMasry et al. (2022)
81.4%自己申告
Vision
ScienceQA Visual
97.5%自己申告
DocVQADocVQA (2020)
93.2%自己申告
MMBench
86.7%自己申告
POPE
85.6%自己申告
OCRBench
84.4%自己申告
AI2D
82.3%自己申告
TextVQA
75.6%自己申告
InfoVQA
72.7%自己申告
BLINK
61.3%自己申告
MMMU-Pro
38.5%自己申告
AA評価指数
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))69.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))48.5
Gpqa(NYU + Cohere + Anthropic (2023))31.5
Livecodebench(UC Berkeley + MIT + Cornell (2024))13.1
Aime(MAA (Mathematical Association of America))9.3
Intelligence Index(Artificial Analysis)5.8
Hle(Center for AI Safety + Scale AI (2025))5.0
LLM Statsカテゴリスコア
(LLM Stats (zeroeval))Safety90
Image To Text80
Multimodal70
Reasoning70
General70
Vision70
Math60
Spatial Reasoning60
Healthcare60
3d60
価格設定
入力価格無料
出力価格無料
混合価格(3:1)無料
速度
トークン/秒17.2
初トークン遅延0.38s
初回答遅延0.38s
プロバイダー価格ランキング
プロバイダー価格ランキング
2 プロバイダー
最安: Azure Cognitive Services最高: Azure
プロバイダー入力出力
1Azure Cognitive Services最安
$0.08
$0.32
2Azure
$0.08
$0.32
このモデルの異なるAPIプロバイダー間の価格を比較。