Llama 3.2 Instruct 90B (Vision)
MetaLlama開源權重Llama 3.2 · 商用許可
描述
Llama 3.2 90B is a large multimodal language model optimized for visual recognition, image reasoning, and captioning tasks. It supports a context length of 128,000 tokens and is designed for deployment on edge and mobile devices, offering state-of-the-art performance in image understanding and generative tasks.
發布日期
2024-09-25
參數規模
90.0B
上下文長度
—
支援模態
image, text
能力雷達圖
24
general
21
coding
30
reasoning
31
science
27
agents
85
multimodal
排行榜排名
基準測試分數 (LLM Stats)
(LLM Stats (zeroeval))General
MMLU
86.0%自報
Math
MGSM
86.9%自報
MATH
68.0%自報
MathVista
57.3%自報
Multimodal
MMMU
60.3%自報
Reasoning
ChartQAMasry et al. (2022)
85.5%自報
GPQANYU + Cohere + Anthropic (2023)
46.7%自報
Vision
AI2D
92.3%自報
DocVQADocVQA (2020)
90.1%自報
VQAv2
78.1%自報
TextVQA
73.5%自報
InfographicsQA
56.8%自報
MMMU-Pro
45.2%自報
AA 評測指數
(Artificial Analysis)Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))67.1
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))62.9
Gpqa(NYU + Cohere + Anthropic (2023))43.2
Livecodebench(UC Berkeley + MIT + Cornell (2024))21.4
Intelligence Index(Artificial Analysis)6.4
Aime(MAA (Mathematical Association of America))5.0
Hle(Center for AI Safety + Scale AI (2025))4.5
LLM Stats 分類評分
(LLM Stats (zeroeval))Language90
Legal90
Finance90
Image To Text80
Math70
Multimodal70
Reasoning70
General70
Healthcare70
Vision70
Physics50
Biology50
Chemistry50
定價
輸入價格免費
輸出價格免費
混合價格(3:1)免費
速度
Tokens/秒0.0
首Token延遲0.00s
首回答延遲0.00s
供應商價格排行
暫無提供商資料