Phi-4 Mini Instruct
MicrosoftPhiOpen WeightMIT · Commercial OK
Description
Phi 4 Mini Instruct is a lightweight (3.8B parameters) open model built upon synthetic data and filtered web data, focusing on high-quality reasoning. It supports a 128K token context length and is enhanced for instruction adherence and safety via supervised fine-tuning and direct preference optimization.
Release Date
2024-02-26
Parameters
3.8B
Context Length
—
Modalities
—
Capability Radar
18
general
7
coding
18
reasoning
24
science
15
agents
0
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Code Ranking | 592 | 8.0 | AA |
| General Ranking | 611 | 16.0 | AA |
| Science | 594 | 16.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))General
MMLU
67.3%SR
TruthfulQA
66.4%SR
Multilingual MMLU
49.3%SR
Arena Hard
32.8%SR
Language
BoolQ
81.2%SR
MMLU-Pro
52.8%SR
Math
MATH-500
94.6%SR
GSM8k
88.6%SR
MATH
64.0%SR
MGSM
63.9%SR
AIMEMAA
57.5%SR
Reasoning
ARC-C
83.7%SR
OpenBookQA
79.2%SR
PIQA
77.6%SR
Social IQa
72.5%SR
BIG-Bench Hard
70.4%SR
HellaSwagAI2 (2019)
69.1%SR
Winogrande
67.0%SR
GPQANYU + Cohere + Anthropic (2023)
52.0%SR
AA Evaluation Indices
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))69.6
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))46.5
Gpqa(NYU + Cohere + Anthropic (2023))33.1
Ifbench(Google Research (2023))21.1
Lcr(Artificial Analysis)15.3
Livecodebench(UC Berkeley + MIT + Cornell (2024))12.6
Tau2(Sierra + U Toronto + Vector Institute (2025))8.2
Math Index(Artificial Analysis)6.7
Aime 25(MAA (Mathematical Association of America))6.7
Intelligence Index(Artificial Analysis)6.3
Hle(Center for AI Safety + Scale AI (2025))4.4
Coding Index(Artificial Analysis)3.8
Aime(MAA (Mathematical Association of America))3.0
Terminalbench V2 10.4
Terminalbench Hard(Stanford × Laude Institute (2026))0.0
LLM Stats Category Scores
(LLM Stats (zeroeval))Math70
Psychology70
Reasoning70
General70
Language60
Legal60
Finance60
Healthcare60
Physics50
Creativity50
Chat30
Biology30
Chemistry30
Writing30
Pricing
Input PriceFree
Output PriceFree
Blended Price (3:1)Free
Speed
Tokens/sec47.3
Time to First Token0.33s
Time to Answer0.33s
Provider Price Ranking
Provider Price Ranking
2 providers
Cheapest: Azure Cognitive ServicesMost Expensive: Azure
ProviderInputOutput
1Azure Cognitive ServicesCheapest
$0.075
$0.3
2Azure
$0.075
$0.3
Compare pricing across different API providers for this model.