Phi-4
MicrosoftPhiOpen WeightMIT · Commercial OK
Description
phi-4 is a state-of-the-art open model built to excel at advanced reasoning, coding, and knowledge tasks. It leverages a blend of synthetic data, filtered web data, academic texts, and supervised fine-tuning for precision, alignment, and safety.
Release Date
2024-12-12
Parameters
14.7B
Context Length
16K
Modalities
text
Capability Radar
25
general
23
coding
30
reasoning
41
science
28
agents
0
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Code Ranking | 586 | 10.0 | AA |
| General Ranking | 569 | 21.0 | AA |
| Science | 475 | 30.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Chat
IFEvalGoogle Research (2023)
83.4%SR
Factuality
SimpleQA
3.0%SR
General
MMLU
84.8%SR
Arena Hard
73.3%SR
Language
MMLU-Pro
74.3%SR
Math
MGSM
80.6%SR
MATH
80.4%SR
OmniMath
76.6%SR
AIME 2024
75.3%SR
AIME 2025
62.9%SR
LiveBench
47.6%SR
Reasoning
FlenQA
97.7%SR
HumanEval+
92.9%SR
HumanEvalOpenAI (2021)
82.6%SR
DROP
75.5%SR
PhiBench
70.6%SR
GPQANYU + Cohere + Anthropic (2023)
65.8%SR
LiveCodeBench
53.8%SR
AA Evaluation Indices
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))81.0
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))71.4
Gpqa(NYU + Cohere + Anthropic (2023))57.5
Ifbench(Google Research (2023))23.5
Livecodebench(UC Berkeley + MIT + Cornell (2024))23.1
Math Index(Artificial Analysis)18.0
Aime 25(MAA (Mathematical Association of America))18.0
Aime(MAA (Mathematical Association of America))14.3
Intelligence Index(Artificial Analysis)5.9
Hle(Center for AI Safety + Scale AI (2025))3.8
Terminalbench Hard(Stanford × Laude Institute (2026))3.8
Lcr(Artificial Analysis)0.0
Tau2(Sierra + U Toronto + Vector Institute (2025))0.0
LLM Stats Category Scores
(LLM Stats (zeroeval))Language80
Legal80
Finance80
Healthcare80
Code80
Creativity80
Writing80
Chat70
Math70
Reasoning70
General70
Instruction Following60
Physics60
Structured Output60
Biology60
Chemistry60
Factuality0
Pricing
Input Price$0.125 / 1M tokens
Output Price$0.5 / 1M tokens
Blended Price (3:1)$0.219 / 1M tokens
Speed
Tokens/sec43.9
Time to First Token0.98s
Time to Answer0.98s
Provider Price Ranking
Provider Price Ranking
6 providers
Cheapest: DeepInfraMost Expensive: Azure
ProviderInputOutput
1DeepInfraCheapest
$0
$0
2OpenRouter
$0.07
$0.14
3Kilo Gateway
$0.07
$0.14
4MicrosoftPRIMARY
$0.125
$0.5
5Azure Cognitive Services
$0.125
$0.5
6Azure
$0.125
$0.5
Compare pricing across different API providers for this model.