メインコンテンツへスキップ

DeepSeek R1 Zero

DeepSeekDeepSeekオープンウエイトMIT · 商用利用可

説明

DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorporates cold-start data before RL. DeepSeek-R1 achieves performance comparable to OpenAI-o1 across math, code, and reasoning tasks.

リリース日
2025-01-20
パラメータ
671.0B
コンテキスト長
モダリティ

能力レーダー

80
general
50
coding
90
reasoning
60
science推定
78
agents
0
multimodal

専用の科学ベンチマークがない場合、Science は LLM Stats の科学スコアまたは推論能力から推定します。

ランキング

ランキングデータがありません

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Biology

GPQANYU + Cohere + Anthropic (2023)73.3%自己申告

Code

LiveCodeBench50.0%自己申告

Math

MATH-50095.9%自己申告
AIME 202486.7%自己申告

AA評価指数

(Artificial Analysis)

AA評価データがありません

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Math
90
Reasoning
80
General
80
Physics
70
Biology
70
Chemistry
70
Code
50

価格設定

価格データがありません

速度

速度データがありません

プロバイダー価格ランキング

プロバイダーデータがありません

外部リンク