DeepSeek-V2.5 (Dec '24)
DeepSeekDeepSeekオープンウエイトdeepseek
説明
DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It better aligns with human preferences and has been optimized in various aspects, including writing and instruction following.
リリース日
2024-12-10
パラメータ
236.0B
コンテキスト長
164K
モダリティ
text
能力レーダー
28
general
60
coding
63
reasoning
42
science
62
agents
0
multimodal
ランキング
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Chat
MT-Bench
0.90 / 100自己申告
General
MMLU
80.4%自己申告
AlignBench
80.4%自己申告
DS-FIM-Eval
78.3%自己申告
Arena Hard
76.2%自己申告
AlpacaEval 2.0
50.5%自己申告
Math
GSM8k
95.1%自己申告
MATH
74.7%自己申告
Reasoning
HumanEvalOpenAI (2021)
89.0%自己申告
BBH
84.3%自己申告
HumanEval-Mul
73.8%自己申告
Aider
72.2%自己申告
DS-Arena-Code
63.1%自己申告
LiveCodeBench(01-09)
41.8%自己申告
SWE-Bench Verified
16.8%自己申告
AA評価指数
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))76.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))66.6
Gpqa(NYU + Cohere + Anthropic (2023))42.3
Intelligence Index(Artificial Analysis)6.6
LLM Statsカテゴリスコア
(LLM Stats (zeroeval))Roleplay90
Communication90
Chat80
Language80
Legal80
Math80
Finance80
Healthcare80
Reasoning70
General70
Creativity70
Writing70
Code60
Frontend Development20
価格設定
入力価格無料
出力価格無料
混合価格(3:1)無料
速度
トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s
プロバイダー価格ランキング
プロバイダーデータがありません