メインコンテンツへスキップ

Phi-3.5-vision-instruct

MicrosoftPhiオープンウエイトMIT · 商用利用可

説明

Phi-3.5-vision-instruct is a 4.2B-parameter open multimodal model with up to 128K context tokens. It emphasizes multi-frame image understanding and reasoning, boosting performance on single-image benchmarks while enabling multi-image comparison, summarization, and even video analysis. The model underwent safety post-training for improved instruction-following, alignment, and robust handling of visual and text inputs, and is released under the MIT license.

リリース日
2024-08-23
パラメータ
4.2B
コンテキスト長
モダリティ

能力レーダー

70
general
0
coding
40
reasoning
34
science推定
28
agents
70
multimodal

専用の科学ベンチマークがない場合、Science は LLM Stats の科学スコアまたは推論能力から推定します。

ランキング

ドメイン#順位スコアソース
マルチモーダルランキング66
37.0
LS

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

General

MMMU43.0%自己申告

Image To Text

TextVQA72.0%自己申告

Math

ScienceQA91.3%自己申告
MathVista43.9%自己申告
InterGPS36.3%自己申告

Multimodal

POPE86.1%自己申告
MMBench81.9%自己申告
ChartQAMasry et al. (2022)81.8%自己申告
AI2D78.1%自己申告

AA評価指数

(Artificial Analysis)

AA評価データがありません

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Multimodal
70
Reasoning
70
Image To Text
70
General
70
Vision
70
Math
40
Healthcare
40

価格設定

価格データがありません

速度

速度データがありません

プロバイダー価格ランキング

プロバイダーデータがありません

外部リンク