Skip to main content

DeepSeek VL2 Small

DeepSeekDeepSeekOpen Weightdeepseek

Description

An advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.

Release Date
2024-12-13
Parameters
16.0B
Context Length
164K
Modalities
text

Capability Radar

70
general
0
coding
60
reasoning
43
scienceest.
42
agents
0
multimodal

Science is estimated from LLM Stats science scores or reasoning when no dedicated science benchmarks are available.

Rankings

Domain#RankScoreSource
Multimodal Ranking143
22.0
LS

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Math

MathVista60.7%SR

Multimodal

MMMU48.0%SR

Reasoning

ChartQAMasry et al. (2022)84.5%SR

Vision

DocVQADocVQA (2020)92.3%SR
TextVQA83.4%SR
OCRBench83.4%SR
MMBench80.3%SR
AI2D80.0%SR
MMBench-V1.179.3%SR
InfoVQA75.8%SR
RealWorldQA65.4%SR
MMT-Bench62.9%SR
MMStar57.0%SR
MME21.2%SR

AA Evaluation Indices

(Artificial Analysis)

No AA evaluation data available

LLM Stats Category Scores

(LLM Stats (zeroeval))
Image To Text
90
Multimodal
70
Spatial Reasoning
70
General
70
Vision
70
Math
60
Reasoning
60
Healthcare
50

Pricing

Input Price$0.2574 / 1M tokens
Output Price$1.0287 / 1M tokens
Blended Price (3:1)$0.45023 / 1M tokens

Speed

No speed data available

Provider Price Ranking

Provider Price Ranking

1 providers

ProviderInputOutput
1DeepSeekPRIMARY
$0.2574
$1.0287

Compare pricing across different API providers for this model.

External Sources