DeepSeek-V2.5 (Dec '24)
DeepSeekDeepSeekオープンウエイトdeepseek
説明
DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It better aligns with human preferences and has been optimized in various aspects, including writing and instruction following.
リリース日
2024-12-10
パラメータ
236.0B
コンテキスト長
—
モダリティ
text
能力レーダー
7
general
60
coding
76
reasoning
68
science推定
71
agents
0
multimodal
専用の科学ベンチマークがない場合、Science は LLM Stats の科学スコアまたは推論能力から推定します。
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| 総合ランキング | 560 | 9.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Code
HumanEvalOpenAI (2021)
89.0%自己申告
Aider
72.2%自己申告
SWE-Bench Verified
16.8%自己申告
Communication
MT-Bench
0.90 / 100自己申告
Creativity
AlignBench
80.4%自己申告
Arena Hard
76.2%自己申告
AlpacaEval 2.0
50.5%自己申告
Finance
MMLU
80.4%自己申告
General
DS-FIM-Eval
78.3%自己申告
LiveCodeBench(01-09)
41.8%自己申告
Language
BBH
84.3%自己申告
Math
GSM8k
95.1%自己申告
MATH
74.7%自己申告
Reasoning
HumanEval-Mul
73.8%自己申告
DS-Arena-Code
63.1%自己申告
AA評価指数
(Artificial Analysis)Intelligence Index(Artificial Analysis)6.5
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))0.8
LLM Statsカテゴリスコア
(LLM Stats (zeroeval))Roleplay90
Communication90
Legal80
Math80
Language80
Finance80
Healthcare80
Reasoning70
General70
Creativity70
Writing70
Code60
Frontend Development20
価格設定
入力価格無料
出力価格無料
混合価格(3:1)無料
速度
トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s
プロバイダー価格ランキング
プロバイダーデータがありません