Phi-3.5-vision-instruct
MicrosoftPhiОткрытые весаMIT · Коммерческое использование
Описание
Phi-3.5-vision-instruct is a 4.2B-parameter open multimodal model with up to 128K context tokens. It emphasizes multi-frame image understanding and reasoning, boosting performance on single-image benchmarks while enabling multi-image comparison, summarization, and even video analysis. The model underwent safety post-training for improved instruction-following, alignment, and robust handling of visual and text inputs, and is released under the MIT license.
Дата выхода
2024-08-23
Параметры
4.2B
Длина контекста
—
Модальности
—
Радар способностей
70
general
0
coding
40
reasoning
34
scienceоцен.
28
agents
70
multimodal
При отсутствии специализированных научных бенчмарков Science оценивается по научным категориям LLM Stats или на основе рассуждений.
Рейтинги
| Домен | #Место | Оценка | Источник |
|---|---|---|---|
| Мультимодальный рейтинг | 69 | 37.0 | LS |
Оценки бенчмарков (LLM Stats)
(LLM Stats (zeroeval))General
MMMU
43.0%Сам.
Image To Text
TextVQA
72.0%Сам.
Math
ScienceQA
91.3%Сам.
MathVista
43.9%Сам.
InterGPS
36.3%Сам.
Multimodal
POPE
86.1%Сам.
MMBench
81.9%Сам.
ChartQAMasry et al. (2022)
81.8%Сам.
AI2D
78.1%Сам.
Индексы оценки AA
(Artificial Analysis)Нет данных AA оценки
Оценки категорий LLM Stats
(LLM Stats (zeroeval))Safety90
Multimodal70
Reasoning70
Image To Text70
General70
Vision70
Math40
Healthcare40
Цены
Нет данных о ценах
Скорость
Нет данных о скорости
Рейтинг цен провайдеров
Нет данных провайдеров