Phi-3.5-MoE-instruct
MicrosoftPhiОткрытые весаMIT · Коммерческое использование
Описание
Phi-3.5-MoE-instruct is a mixture-of-experts model with ~42B total parameters (6.6B active) and a 128K context window. It excels at reasoning, math, coding, and multilingual tasks, outperforming larger dense models in many benchmarks. It underwent a thorough safety post-training process (SFT + DPO) and is licensed under MIT. This model is ideal for scenarios where efficiency and high performance are both required, particularly in multi-lingual or reasoning-intensive tasks.
Дата выхода
2024-08-23
Параметры
60.0B
Длина контекста
—
Модальности
—
Радар способностей
70
general
70
coding
70
reasoning
34
scienceоцен.
70
agents
0
multimodal
При отсутствии специализированных научных бенчмарков Science оценивается по научным категориям LLM Stats или на основе рассуждений.
Рейтинги
Нет данных рейтинга
Оценки бенчмарков (LLM Stats)
(LLM Stats (zeroeval))General
MMLU
78.9%Сам.
TruthfulQA
77.5%Сам.
Arena Hard
37.9%Сам.
Language
BoolQ
84.6%Сам.
MMMLU
69.9%Сам.
MEGA TyDi QA
67.1%Сам.
MEGA MLQA
65.3%Сам.
MEGA UDPOS
60.4%Сам.
MMLU-Pro
45.3%Сам.
Long Context
RULER
87.1%Сам.
RepoQA
85.0%Сам.
Math
GSM8k
88.7%Сам.
MATH
59.5%Сам.
MGSM
58.7%Сам.
Reasoning
ARC-C
91.0%Сам.
OpenBookQA
89.6%Сам.
PIQA
88.6%Сам.
HellaSwagAI2 (2019)
83.8%Сам.
MEGA XStoryCloze
82.8%Сам.
Winogrande
81.3%Сам.
MBPP
0.81 / 100Сам.
BIG-Bench Hard
79.1%Сам.
Social IQa
78.0%Сам.
MEGA XCOPA
76.6%Сам.
HumanEvalOpenAI (2021)
70.7%Сам.
Qasper
40.0%Сам.
GPQANYU + Cohere + Anthropic (2023)
36.8%Сам.
Summarization
GovReport
26.4%Сам.
SQuALITY
24.1%Сам.
QMSum
19.9%Сам.
SummScreenFD
16.9%Сам.
Индексы оценки AA
(Artificial Analysis)Нет данных AA оценки
Оценки категорий LLM Stats
(LLM Stats (zeroeval))Psychology80
Language70
Legal70
Math70
Reasoning70
Finance70
General70
Healthcare70
Code70
Long Context60
Physics60
Creativity60
Chat40
Biology40
Chemistry40
Writing40
Summarization20
Цены
Нет данных о ценах
Скорость
Нет данных о скорости
Рейтинг цен провайдеров
Нет данных провайдеров