Llama 3.2 Instruct 11B (Vision)
MetaLlamaオープンウエイトLlama 3.2 Community License
説明
Llama 3.2 11B Vision Instruct is an instruction-tuned multimodal large language model optimized for visual recognition, image reasoning, captioning, and answering general questions about an image. It accepts text and images as input and generates text as output.
リリース日
2024-09-25
パラメータ
10.6B
コンテキスト長
—
モダリティ
image, text
能力レーダー
18
general
11
coding
13
reasoning
17
science
12
agents
90
multimodal
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| コーディングランキング | 589 | 9.0 | AA |
| 総合ランキング | 580 | 20.0 | AA |
| マルチモーダルランキング | 123 | 37.0 | LS |
| 科学 | 641 | 11.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))General
MMLU
73.0%自己申告
Math
MGSM
68.9%自己申告
MATH
51.9%自己申告
MathVista
51.5%自己申告
Multimodal
MMMU
50.7%自己申告
Reasoning
ChartQAMasry et al. (2022)
83.4%自己申告
GPQANYU + Cohere + Anthropic (2023)
32.8%自己申告
Vision
AI2D
91.1%自己申告
DocVQADocVQA (2020)
88.4%自己申告
VQAv2 (test)
75.2%自己申告
MMMU-Pro
33.0%自己申告
AA評価指数
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))51.6
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))46.4
Ifbench(Google Research (2023))30.4
Gpqa(NYU + Cohere + Anthropic (2023))22.1
Tau2(Sierra + U Toronto + Vector Institute (2025))14.6
Lcr(Artificial Analysis)12.7
Livecodebench(UC Berkeley + MIT + Cornell (2024))11.0
Aime(MAA (Mathematical Association of America))9.3
Hle(Center for AI Safety + Scale AI (2025))5.5
Intelligence Index(Artificial Analysis)5.4
Math Index(Artificial Analysis)1.7
Aime 25(MAA (Mathematical Association of America))1.7
Terminalbench Hard(Stanford × Laude Institute (2026))0.8
LLM Statsカテゴリスコア
(LLM Stats (zeroeval))Image To Text90
Language70
Legal70
Multimodal70
Finance70
Vision70
Math60
Reasoning60
General60
Healthcare60
Physics30
Biology30
Chemistry30
価格設定
入力価格$0.345 / 1Mトークン
出力価格$0.345 / 1Mトークン
混合価格(3:1)$0.345 / 1Mトークン
速度
トークン/秒13.8
初トークン遅延1.28s
初回答遅延1.28s
プロバイダー価格ランキング
プロバイダー価格ランキング
1 プロバイダー
プロバイダー入力出力
1Metaプライマリ
$0.345
$0.345
このモデルの異なるAPIプロバイダー間の価格を比較。