Phi-4 Mini Instruct
MicrosoftPhiОткрытые весаMIT · Коммерческое использование
Описание
Phi 4 Mini Instruct is a lightweight (3.8B parameters) open model built upon synthetic data and filtered web data, focusing on high-quality reasoning. It supports a 128K token context length and is enhanced for instruction adherence and safety via supervised fine-tuning and direct preference optimization.
Дата выхода
2024-02-26
Параметры
3.8B
Длина контекста
—
Модальности
—
Радар способностей
18
general
7
coding
18
reasoning
24
science
15
agents
0
multimodal
Рейтинги
| Домен | #Место | Оценка | Источник |
|---|---|---|---|
| Рейтинг кодинга | 599 | 8.0 | AA |
| Общий рейтинг | 616 | 16.0 | AA |
| Наука | 601 | 16.0 | AA |
Оценки бенчмарков (LLM Stats)
(LLM Stats (zeroeval))General
MMLU
67.3%Сам.
TruthfulQA
66.4%Сам.
Multilingual MMLU
49.3%Сам.
Arena Hard
32.8%Сам.
Language
BoolQ
81.2%Сам.
MMLU-Pro
52.8%Сам.
Math
MATH-500
94.6%Сам.
GSM8k
88.6%Сам.
MATH
64.0%Сам.
MGSM
63.9%Сам.
AIMEMAA
57.5%Сам.
Reasoning
ARC-C
83.7%Сам.
OpenBookQA
79.2%Сам.
PIQA
77.6%Сам.
Social IQa
72.5%Сам.
BIG-Bench Hard
70.4%Сам.
HellaSwagAI2 (2019)
69.1%Сам.
Winogrande
67.0%Сам.
GPQANYU + Cohere + Anthropic (2023)
52.0%Сам.
Индексы оценки AA
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))69.6
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))46.5
Gpqa(NYU + Cohere + Anthropic (2023))33.1
Ifbench(Google Research (2023))21.1
Lcr(Artificial Analysis)15.3
Livecodebench(UC Berkeley + MIT + Cornell (2024))12.6
Tau2(Sierra + U Toronto + Vector Institute (2025))8.2
Math Index(Artificial Analysis)6.7
Aime 25(MAA (Mathematical Association of America))6.7
Intelligence Index(Artificial Analysis)6.3
Hle(Center for AI Safety + Scale AI (2025))4.4
Coding Index(Artificial Analysis)3.8
Aime(MAA (Mathematical Association of America))3.0
Terminalbench V2 10.4
Terminalbench Hard(Stanford × Laude Institute (2026))0.0
Оценки категорий LLM Stats
(LLM Stats (zeroeval))Math70
Psychology70
Reasoning70
General70
Language60
Legal60
Finance60
Healthcare60
Physics50
Creativity50
Chat30
Biology30
Chemistry30
Writing30
Цены
Цена вводаБесплатно
Цена выводаБесплатно
Смешанная цена (3:1)Бесплатно
Скорость
Токенов/сек46.0
Задержка первого токена0.33s
Время до первого ответа0.33s
Рейтинг цен провайдеров
Рейтинг цен провайдеров
2 провайдеров
Самый дешевый: Azure Cognitive ServicesСамый дорогой: Azure
ПровайдерВводВывод
1Azure Cognitive ServicesСамый дешевый
$0.075
$0.3
2Azure
$0.075
$0.3
Сравнение цен разных API-провайдеров для этой модели.