मुख्य सामग्री पर जाएं

Phi-4 Multimodal Instruct

MicrosoftPhiओपन वेटMIT · व्यावसायिक उपयोग

विवरण

Phi-4-multimodal-instruct is a lightweight (5.57B parameters) open multimodal foundation model that leverages research and datasets from Phi-3.5 and 4.0. It processes text, image, and audio inputs to generate text outputs, supporting a 128K token context length. Enhanced via SFT, DPO, and RLHF for instruction following and safety.

रिलीज़ तिथि
2025-02-26
पैरामीटर
5.6B
संदर्भ लंबाई
—
मोडैलिटीज़
image, text

क्षमता रडार

18
general
13
coding
32
reasoning
23
science
26
agents
85
multimodal

रैंकिंग

बेंचमार्क स्कोर (LLM Stats)

(LLM Stats (zeroeval))

Math

MathVista62.4%स्वयं
InterGPS48.6%स्वयं

Multimodal

MMMU55.1%स्वयं
Video-MME55.0%स्वयं

Reasoning

ChartQAMasry et al. (2022)81.4%स्वयं

Vision

ScienceQA Visual97.5%स्वयं
DocVQADocVQA (2020)93.2%स्वयं
MMBench86.7%स्वयं
POPE85.6%स्वयं
OCRBench84.4%स्वयं
AI2D82.3%स्वयं
TextVQA75.6%स्वयं
InfoVQA72.7%स्वयं
BLINK61.3%स्वयं
MMMU-Pro38.5%स्वयं

AA मूल्यांकन सूचकांक

(Artificial Analysis)
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
69.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
48.5
Gpqa(NYU + Cohere + Anthropic (2023))
31.5
Livecodebench(UC Berkeley + MIT + Cornell (2024))
13.1
Aime(MAA (Mathematical Association of America))
9.3
Intelligence Index(Artificial Analysis)
5.8
Hle(Center for AI Safety + Scale AI (2025))
5.0

LLM Stats श्रेणी स्कोर

(LLM Stats (zeroeval))
Safety
90
Image To Text
80
Multimodal
70
Reasoning
70
General
70
Vision
70
Math
60
Spatial Reasoning
60
Healthcare
60
3d
60

मूल्य निर्धारण

इनपुट मूल्यमुफ्त
आउटपुट मूल्यमुफ्त
मिश्रित मूल्य (3:1)मुफ्त

गति

टोकन/सेकंड17.2
पहले टोकन में देरी0.36s
पहले उत्तर में देरी0.36s

प्रदाता मूल्य रैंकिंग

प्रदाता मूल्य रैंकिंग

2 प्रदाता

सबसे सस्ता: Azure Cognitive Servicesसबसे महंगा: Azure
प्रदाताइनपुटआउटपुट
1Azure Cognitive Servicesसबसे सस्ता
$0.08
$0.32
2Azure
$0.08
$0.32

इस मॉडल के लिए विभिन्न API प्रदाताओं के मूल्य निर्धारण की तुलना करें।

बाहरी लिंक