DeepSeek-V2.5 (Dec '24)
DeepSeekDeepSeekOpen Weightdeepseek
Descripción
DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It better aligns with human preferences and has been optimized in various aspects, including writing and instruction following.
Fecha de lanzamiento
2024-12-10
Parámetros
236.0B
Longitud del contexto
164K
Modalidades
text
Radar de capacidades
28
general
60
coding
63
reasoning
42
science
62
agents
0
multimodal
Rankings
| Dominio | #Posición | Puntuación | Fuente |
|---|---|---|---|
| Ranking general | 491 | 29.0 | AA |
| Ciencia | 398 | 38.0 | AA |
Puntuaciones de benchmarks (LLM Stats)
(LLM Stats (zeroeval))Chat
MT-Bench
0.90 / 100Aut.
General
MMLU
80.4%Aut.
AlignBench
80.4%Aut.
DS-FIM-Eval
78.3%Aut.
Arena Hard
76.2%Aut.
AlpacaEval 2.0
50.5%Aut.
Math
GSM8k
95.1%Aut.
MATH
74.7%Aut.
Reasoning
HumanEvalOpenAI (2021)
89.0%Aut.
BBH
84.3%Aut.
HumanEval-Mul
73.8%Aut.
Aider
72.2%Aut.
DS-Arena-Code
63.1%Aut.
LiveCodeBench(01-09)
41.8%Aut.
SWE-Bench Verified
16.8%Aut.
Índices de evaluación AA
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))76.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))66.6
Gpqa(NYU + Cohere + Anthropic (2023))42.3
Intelligence Index(Artificial Analysis)6.6
Puntuaciones por categoría LLM Stats
(LLM Stats (zeroeval))Roleplay90
Communication90
Chat80
Language80
Legal80
Math80
Finance80
Healthcare80
Reasoning70
General70
Creativity70
Writing70
Code60
Frontend Development20
Precios
Precio de entradaGratis
Precio de salidaGratis
Precio mixto (3:1)Gratis
Velocidad
Tokens/seg0.0
Retraso del primer token0.00s
Tiempo hasta la respuesta0.00s
Ranking de Precios por Proveedor
No hay datos de proveedores disponibles