メインコンテンツへスキップ

Phi-3.5-vision-instruct

MicrosoftPhiオープンウエイトMIT · 商用利用可

説明

Phi-3.5-vision-instruct is a 4.2B-parameter open multimodal model with up to 128K context tokens. It emphasizes multi-frame image understanding and reasoning, boosting performance on single-image benchmarks while enabling multi-image comparison, summarization, and even video analysis. The model underwent safety post-training for improved instruction-following, alignment, and robust handling of visual and text inputs, and is released under the MIT license.

リリース日
2024-08-23
パラメータ
4.2B
コンテキスト長
—
モダリティ
—

能力レーダー

70
general
0
coding
40
reasoning
34
science推定
28
agents
70
multimodal

専用の科学ベンチマークがない場合、Science は LLM Stats の科学スコアまたは推論能力から推定します。

ランキング

ドメイン#順位スコアソース
マルチモーダルランキング135
29.0
LS

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Math

MathVista43.9%自己申告
InterGPS36.3%自己申告

Multimodal

MMMU43.0%自己申告

Reasoning

ScienceQA91.3%自己申告
ChartQAMasry et al. (2022)81.8%自己申告

Vision

POPE86.1%自己申告
MMBench81.9%自己申告
AI2D78.1%自己申告
TextVQA72.0%自己申告

AA評価指数

(Artificial Analysis)

AA評価データがありません

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Safety
90
Image To Text
70
Multimodal
70
Reasoning
70
General
70
Vision
70
Math
40
Healthcare
40

価格設定

価格データがありません

速度

速度データがありません

プロバイダー価格ランキング

プロバイダーデータがありません

外部リンク