DeepSeek VL2 Small
DeepSeekDeepSeekOpen Weightdeepseek
Description
An advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.
Release Date
2024-12-13
Parameters
16.0B
Context Length
164K
Modalities
text
Capability Radar
70
general
0
coding
60
reasoning
43
scienceest.
42
agents
0
multimodal
Science is estimated from LLM Stats science scores or reasoning when no dedicated science benchmarks are available.
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Multimodal Ranking | 143 | 22.0 | LS |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Math
MathVista
60.7%SR
Multimodal
MMMU
48.0%SR
Reasoning
ChartQAMasry et al. (2022)
84.5%SR
Vision
DocVQADocVQA (2020)
92.3%SR
TextVQA
83.4%SR
OCRBench
83.4%SR
MMBench
80.3%SR
AI2D
80.0%SR
MMBench-V1.1
79.3%SR
InfoVQA
75.8%SR
RealWorldQA
65.4%SR
MMT-Bench
62.9%SR
MMStar
57.0%SR
MME
21.2%SR
AA Evaluation Indices
(Artificial Analysis)No AA evaluation data available
LLM Stats Category Scores
(LLM Stats (zeroeval))Image To Text90
Multimodal70
Spatial Reasoning70
General70
Vision70
Math60
Reasoning60
Healthcare50
Pricing
Input Price$0.2574 / 1M tokens
Output Price$1.0287 / 1M tokens
Blended Price (3:1)$0.45023 / 1M tokens
Speed
No speed data available
Provider Price Ranking
Provider Price Ranking
1 providers
ProviderInputOutput
1DeepSeekPRIMARY
$0.2574
$1.0287
Compare pricing across different API providers for this model.