メインコンテンツへスキップ

Kimi K2

KimiKimiオープンウエイトMIT · 商用利用可

説明

Kimi K2-Instruct-0905 is the latest, most capable version of Kimi K2, achieving state-of-the-art performance in frontier knowledge, math, and coding among non-thinking models. This Mixture-of-Experts model features 32 billion activated parameters and 1 trillion total parameters, meticulously optimized for agentic tasks. Key features include enhanced agentic coding intelligence, extended context length to 256K tokens, and a hybrid architecture trained with MuonClip optimizer on 15.5T tokens. The model achieves 65.8% on SWE-bench Verified (single attempt), 47.3% on SWE-bench Multilingual, and excels at tool use with 70.6% on Tau2-retail. It is a reflex-grade model without long thinking, designed to act and execute complex tasks seamlessly.

リリース日
2025-07-11
パラメータ
1.0T
コンテキスト長
131K
モダリティ
text

能力レーダー

37
general
51
coding
68
reasoning
48
science
60
agents
0
multimodal

ランキング

ドメイン#順位スコアソース
エージェント能力102
39.0
LS
コーディングランキング231
50.0
AA
総合ランキング240
51.0
AA
科学255
49.0
AA

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Agents

Terminal-Bench30.0%自己申告
Terminus25.0%自己申告

Biology

GPQANYU + Cohere + Anthropic (2023)75.1%自己申告

Chemistry

SuperGPQA57.2%自己申告

Code

HumanEvalOpenAI (2021)93.3%自己申告
EvalPlus0.80 / 100自己申告
SWE-bench Verified (Agentic Coding)65.8%自己申告
SWE-Bench Verified65.8%自己申告
Aider-Polyglot60.0%自己申告
LiveCodeBench53.7%自己申告
SWE-bench Multilingual47.3%自己申告

Communication

Tau2 Retail70.6%自己申告
Tau2 Telecom65.8%自己申告
Tau2 Airline56.5%自己申告
Multi-Challenge54.1%自己申告

Factuality

SimpleQA31.0%自己申告

Finance

MMLU89.5%自己申告
MMLU-Pro81.1%自己申告
ACEBench76.5%自己申告

General

MMLU-Redux92.7%自己申告
C-Eval92.5%自己申告
MMLU-redux-2.090.2%自己申告
IFEvalGoogle Research (2023)89.8%自己申告
MultiPL-E85.7%自己申告
TriviaQA85.1%自己申告
CSimpleQA78.4%自己申告
LiveBench76.4%自己申告
LiveCodeBench v653.7%自己申告
SWE-bench Verified (Agentless)51.8%自己申告

Math

MATH-50097.4%自己申告
GSM8k97.3%自己申告
CBNSL95.6%自己申告
CNMO 202474.3%自己申告
MATH70.2%自己申告
AIME 202469.6%自己申告
PolyMath-en65.1%自己申告
AIME 202549.5%自己申告
HMMT 202538.8%自己申告
Humanity's Last Exam4.7%自己申告

Reasoning

AutoLogi89.5%自己申告
ZebraLogic89.0%自己申告
HumanEval-ER81.1%自己申告
MuSR76.4%自己申告
SWE-bench Verified (Multiple Attempts)71.6%自己申告
OJBench27.1%自己申告

AA評価指数

(Artificial Analysis)
Math Index(Artificial Analysis)
57.0
Intelligence Index(Artificial Analysis)
19.7
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
1.0
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
0.8
Gpqa(NYU + Cohere + Anthropic (2023))
0.8
Aime(MAA (Mathematical Association of America))
0.7
Tau2(Sierra + U Toronto + Vector Institute (2025))
0.6
Aime 25(MAA (Mathematical Association of America))
0.6
Livecodebench(UC Berkeley + MIT + Cornell (2024))
0.6
Lcr(Artificial Analysis)
0.5
Ifbench(Google Research (2023))
0.4
Scicode(UIUC + Argonne National Lab (2024))
0.3
Terminalbench Hard(Stanford × Laude Institute (2026))
0.2
Hle(Center for AI Safety + Scale AI (2025))
0.1

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Structured Output
90
Instruction Following
90
Language
90
Legal
80
Finance
80
Healthcare
80
Biology
80
Math
70
Physics
70
Frontend Development
70
Chemistry
70
Reasoning
60
General
60
Communication
60
Economics
60
Tool Calling
60
Code
50
Factuality
30
Agents
20
Vision
0

価格設定

入力価格$0.57 / 1Mトークン
出力価格$2.3 / 1Mトークン
混合価格(3:1)$1.002 / 1Mトークン

速度

トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s

プロバイダー価格ランキング

プロバイダー価格ランキング

7 プロバイダー

最安: OpenCode Zen最高: LLM Gateway
プロバイダー入力出力
1OpenCode Zen最安
$0.4
$2.5
2FastRouter
$0.55
$2.2
3Kimiプライマリ
$0.57
$2.3
4OpenRouter
$0.57
$2.3
5Kilo Gateway
$0.57
$2.3
6Vercel AI Gateway
$0.57
$2.3
7LLM Gateway
$0.57
$2.3

このモデルの異なるAPIプロバイダー間の価格を比較。

外部リンク