Phi-4 Multimodal Instruct
MicrosoftPhiOpen WeightMIT · Commercial OK
Description
Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.
Release Date
2025-02-26
Parameters
5.6B
Context Length
—
Modalities
image, text
Capability Radar
18
general
13
coding
32
reasoning
23
science
26
agents
85
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Code Ranking | 554 | 14.0 | AA |
| General Ranking | 577 | 20.0 | AA |
| Multimodal Ranking | 140 | 23.0 | LS |
| Science | 598 | 16.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Math
MathVista
62.4%SR
InterGPS
48.6%SR
Multimodal
MMMU
55.1%SR
Video-MME
55.0%SR
Reasoning
ChartQAMasry et al. (2022)
81.4%SR
Vision
ScienceQA Visual
97.5%SR
DocVQADocVQA (2020)
93.2%SR
MMBench
86.7%SR
POPE
85.6%SR
OCRBench
84.4%SR
AI2D
82.3%SR
TextVQA
75.6%SR
InfoVQA
72.7%SR
BLINK
61.3%SR
MMMU-Pro
38.5%SR
AA Evaluation Indices
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))69.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))48.5
Gpqa(NYU + Cohere + Anthropic (2023))31.5
Livecodebench(UC Berkeley + MIT + Cornell (2024))13.1
Aime(MAA (Mathematical Association of America))9.3
Intelligence Index(Artificial Analysis)5.8
Hle(Center for AI Safety + Scale AI (2025))5.0
LLM Stats Category Scores
(LLM Stats (zeroeval))Safety90
Image To Text80
Multimodal70
Reasoning70
General70
Vision70
Math60
Spatial Reasoning60
Healthcare60
3d60
Pricing
Input PriceFree
Output PriceFree
Blended Price (3:1)Free
Speed
Tokens/sec17.2
Time to First Token0.36s
Time to Answer0.36s
Provider Price Ranking
Provider Price Ranking
2 providers
Cheapest: Azure Cognitive ServicesMost Expensive: Azure
ProviderInputOutput
1Azure Cognitive ServicesCheapest
$0.08
$0.32
2Azure
$0.08
$0.32
Compare pricing across different API providers for this model.