Phi-4
MicrosoftPhiОткрытые весаMIT · Коммерческое использование
Описание
phi-4 is a state-of-the-art open model built to excel at advanced reasoning, coding, and knowledge tasks. It leverages a blend of synthetic data, filtered web data, academic texts, and supervised fine-tuning for precision, alignment, and safety.
Дата выхода
2024-12-12
Параметры
14.7B
Длина контекста
16K
Модальности
text
Радар способностей
25
general
23
coding
30
reasoning
41
science
28
agents
0
multimodal
Рейтинги
| Домен | #Место | Оценка | Источник |
|---|---|---|---|
| Рейтинг кодинга | 589 | 10.0 | AA |
| Общий рейтинг | 572 | 21.0 | AA |
| Наука | 477 | 30.0 | AA |
Оценки бенчмарков (LLM Stats)
(LLM Stats (zeroeval))Chat
IFEvalGoogle Research (2023)
83.4%Сам.
Factuality
SimpleQA
3.0%Сам.
General
MMLU
84.8%Сам.
Arena Hard
73.3%Сам.
Language
MMLU-Pro
74.3%Сам.
Math
MGSM
80.6%Сам.
MATH
80.4%Сам.
OmniMath
76.6%Сам.
AIME 2024
75.3%Сам.
AIME 2025
62.9%Сам.
LiveBench
47.6%Сам.
Reasoning
FlenQA
97.7%Сам.
HumanEval+
92.9%Сам.
HumanEvalOpenAI (2021)
82.6%Сам.
DROP
75.5%Сам.
PhiBench
70.6%Сам.
GPQANYU + Cohere + Anthropic (2023)
65.8%Сам.
LiveCodeBench
53.8%Сам.
Индексы оценки AA
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))81.0
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))71.4
Gpqa(NYU + Cohere + Anthropic (2023))57.5
Ifbench(Google Research (2023))23.5
Livecodebench(UC Berkeley + MIT + Cornell (2024))23.1
Math Index(Artificial Analysis)18.0
Aime 25(MAA (Mathematical Association of America))18.0
Aime(MAA (Mathematical Association of America))14.3
Intelligence Index(Artificial Analysis)5.9
Hle(Center for AI Safety + Scale AI (2025))3.8
Terminalbench Hard(Stanford × Laude Institute (2026))3.8
Lcr(Artificial Analysis)0.0
Tau2(Sierra + U Toronto + Vector Institute (2025))0.0
Оценки категорий LLM Stats
(LLM Stats (zeroeval))Language80
Legal80
Finance80
Healthcare80
Code80
Creativity80
Writing80
Chat70
Math70
Reasoning70
General70
Instruction Following60
Physics60
Structured Output60
Biology60
Chemistry60
Factuality0
Цены
Цена ввода$0.125 / 1M токенов
Цена вывода$0.5 / 1M токенов
Смешанная цена (3:1)$0.219 / 1M токенов
Скорость
Токенов/сек36.2
Задержка первого токена1.02s
Время до первого ответа1.02s
Рейтинг цен провайдеров
Рейтинг цен провайдеров
6 провайдеров
Самый дешевый: DeepInfraСамый дорогой: Azure
ПровайдерВводВывод
1DeepInfraСамый дешевый
$0
$0
2OpenRouter
$0.07
$0.14
3Kilo Gateway
$0.07
$0.14
4MicrosoftОсновной
$0.125
$0.5
5Azure Cognitive Services
$0.125
$0.5
6Azure
$0.125
$0.5
Сравнение цен разных API-провайдеров для этой модели.