メインコンテンツへスキップ

DeepSeek-V2.5 (Dec '24)

DeepSeekDeepSeekオープンウエイトdeepseek

説明

DeepSeek-V2.5 is an upgraded version that combines DeepSeek-V2-Chat and DeepSeek-Coder-V2-Instruct, integrating general and coding abilities. It better aligns with human preferences and has been optimized in various aspects, including writing and instruction following.

リリース日
2024-12-10
パラメータ
236.0B
コンテキスト長
モダリティ
text

能力レーダー

7
general
60
coding
76
reasoning
68
science推定
71
agents
0
multimodal

専用の科学ベンチマークがない場合、Science は LLM Stats の科学スコアまたは推論能力から推定します。

ランキング

ドメイン#順位スコアソース
総合ランキング560
9.0
AA

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Code

HumanEvalOpenAI (2021)89.0%自己申告
Aider72.2%自己申告
SWE-Bench Verified16.8%自己申告

Communication

MT-Bench0.90 / 100自己申告

Creativity

AlignBench80.4%自己申告
Arena Hard76.2%自己申告
AlpacaEval 2.050.5%自己申告

Finance

MMLU80.4%自己申告

General

DS-FIM-Eval78.3%自己申告
LiveCodeBench(01-09)41.8%自己申告

Language

BBH84.3%自己申告

Math

GSM8k95.1%自己申告
MATH74.7%自己申告

Reasoning

HumanEval-Mul73.8%自己申告
DS-Arena-Code63.1%自己申告

AA評価指数

(Artificial Analysis)
Intelligence Index(Artificial Analysis)
6.5
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
0.8

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Roleplay
90
Communication
90
Legal
80
Math
80
Language
80
Finance
80
Healthcare
80
Reasoning
70
General
70
Creativity
70
Writing
70
Code
60
Frontend Development
20

価格設定

入力価格無料
出力価格無料
混合価格(3:1)無料

速度

トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s

プロバイダー価格ランキング

プロバイダーデータがありません

外部リンク