Phi-4
MicrosoftPhiOpen WeightMIT · Commercial OK
Description
phi-4 is a state-of-the-art open model built to excel at advanced reasoning, coding, and knowledge tasks. It leverages a blend of synthetic data, filtered web data, academic texts, and supervised fine-tuning for precision, alignment, and safety.
Release Date
2024-12-12
Parameters
14.7B
Context Length
16K
Modalities
text
Capability Radar
25
general
24
coding
30
reasoning
36
science
28
agents
0
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Code Ranking | 508 | 10.0 | AA |
| General Ranking | 502 | 21.0 | AA |
| Science | 383 | 35.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Biology
GPQANYU + Cohere + Anthropic (2023)
65.8%SR
Code
HumanEvalOpenAI (2021)
82.6%SR
LiveCodeBench
53.8%SR
Creativity
Arena Hard
73.3%SR
Factuality
SimpleQA
3.0%SR
Finance
MMLU
84.8%SR
MMLU-Pro
74.3%SR
General
IFEvalGoogle Research (2023)
83.4%SR
PhiBench
70.6%SR
LiveBench
47.6%SR
Long Context
FlenQA
97.7%SR
Math
MGSM
80.6%SR
MATH
80.4%SR
OmniMath
76.6%SR
DROP
75.5%SR
AIME 2024
75.3%SR
AIME 2025
62.9%SR
Reasoning
HumanEval+
92.9%SR
AA Evaluation Indices
(Artificial Analysis)Math Index(Artificial Analysis)18.0
Intelligence Index(Artificial Analysis)4.6
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))0.8
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))0.7
Gpqa(NYU + Cohere + Anthropic (2023))0.6
Scicode(UIUC + Argonne National Lab (2024))0.3
Ifbench(Google Research (2023))0.2
Livecodebench(UC Berkeley + MIT + Cornell (2024))0.2
Aime 25(MAA (Mathematical Association of America))0.2
Aime(MAA (Mathematical Association of America))0.1
Hle(Center for AI Safety + Scale AI (2025))0.0
Terminalbench Hard(Stanford × Laude Institute (2026))0.0
Lcr(Artificial Analysis)0.0
Tau2(Sierra + U Toronto + Vector Institute (2025))0.0
LLM Stats Category Scores
(LLM Stats (zeroeval))Legal80
Language80
Finance80
Healthcare80
Code80
Creativity80
Writing80
Math70
Reasoning70
General70
Physics60
Structured Output60
Instruction Following60
Biology60
Chemistry60
Factuality0
Pricing
Input Price$0.125 / 1M tokens
Output Price$0.5 / 1M tokens
Blended Price (3:1)$0.219 / 1M tokens
Speed
Tokens/sec40.9
Time to First Token0.57s
Time to Answer0.57s
Provider Price Ranking
Provider Price Ranking
5 providers
Cheapest: OpenRouterMost Expensive: Azure
ProviderInputOutput
1OpenRouterCheapest
$0.07
$0.14
2Kilo Gateway
$0.07
$0.14
3MicrosoftPRIMARY
$0.125
$0.5
4Azure Cognitive Services
$0.125
$0.5
5Azure
$0.125
$0.5
Compare pricing across different API providers for this model.