DeepSeek-V2.5 (Dec '24)
DeepSeekDeepSeekOpen Weightdeepseek
Description
DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It better aligns with human preferences and has been optimized in various aspects, including writing and instruction following.
Date de sortie
2024-12-10
Paramètres
236.0B
Longueur du contexte
164K
Modalités
text
Radar de capacités
28
general
60
coding
63
reasoning
42
science
62
agents
0
multimodal
Classements
| Domaine | #Rang | Score | Source |
|---|---|---|---|
| Classement général | 488 | 29.0 | AA |
| Science | 395 | 38.0 | AA |
Scores de benchmarks (LLM Stats)
(LLM Stats (zeroeval))Chat
MT-Bench
0.90 / 100Aut.
General
MMLU
80.4%Aut.
AlignBench
80.4%Aut.
DS-FIM-Eval
78.3%Aut.
Arena Hard
76.2%Aut.
AlpacaEval 2.0
50.5%Aut.
Math
GSM8k
95.1%Aut.
MATH
74.7%Aut.
Reasoning
HumanEvalOpenAI (2021)
89.0%Aut.
BBH
84.3%Aut.
HumanEval-Mul
73.8%Aut.
Aider
72.2%Aut.
DS-Arena-Code
63.1%Aut.
LiveCodeBench(01-09)
41.8%Aut.
SWE-Bench Verified
16.8%Aut.
Indices d'évaluation AA
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))76.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))66.6
Gpqa(NYU + Cohere + Anthropic (2023))42.3
Intelligence Index(Artificial Analysis)6.6
Scores par catégorie LLM Stats
(LLM Stats (zeroeval))Roleplay90
Communication90
Chat80
Language80
Legal80
Math80
Finance80
Healthcare80
Reasoning70
General70
Creativity70
Writing70
Code60
Frontend Development20
Tarification
Prix d'entréeGratuit
Prix de sortieGratuit
Prix mixte (3:1)Gratuit
Vitesse
Tokens/sec0.0
Délai du premier token0.00s
Temps de réponse0.00s
Classement des Prix par Fournisseur
Aucune donnée de fournisseur disponible