DeepSeek-V2.5 (Dec '24)
DeepSeekDeepSeekOpen Weightdeepseek
Description
DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It better aligns with human preferences and has been optimized in various aspects, including writing and instruction following.
Release Date
2024-12-10
Parameters
236.0B
Context Length
164K
Modalities
text
Capability Radar
28
general
60
coding
63
reasoning
42
science
62
agents
0
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| General Ranking | 487 | 29.0 | AA |
| Science | 394 | 38.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Chat
MT-Bench
0.90 / 100SR
General
MMLU
80.4%SR
AlignBench
80.4%SR
DS-FIM-Eval
78.3%SR
Arena Hard
76.2%SR
AlpacaEval 2.0
50.5%SR
Math
GSM8k
95.1%SR
MATH
74.7%SR
Reasoning
HumanEvalOpenAI (2021)
89.0%SR
BBH
84.3%SR
HumanEval-Mul
73.8%SR
Aider
72.2%SR
DS-Arena-Code
63.1%SR
LiveCodeBench(01-09)
41.8%SR
SWE-Bench Verified
16.8%SR
AA Evaluation Indices
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))76.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))66.6
Gpqa(NYU + Cohere + Anthropic (2023))42.3
Intelligence Index(Artificial Analysis)6.6
LLM Stats Category Scores
(LLM Stats (zeroeval))Roleplay90
Communication90
Chat80
Language80
Legal80
Math80
Finance80
Healthcare80
Reasoning70
General70
Creativity70
Writing70
Code60
Frontend Development20
Pricing
Input PriceFree
Output PriceFree
Blended Price (3:1)Free
Speed
Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s
Provider Price Ranking
No provider data available