GPT-5.1 (high)
OpenAIGPTProprietary
説明
The best model for coding and agentic tasks with configurable reasoning effort. GPT-5.1 is our flagship model for coding and agentic tasks with configurable reasoning and non-reasoning effort.
リリース日
2025-11-13
パラメータ
—
コンテキスト長
400K
モダリティ
image, text
能力レーダー
44
general
64
coding
93
reasoning
69
science
80
agents
90
multimodal
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| コーディングランキング | 145 | 75.0 | AA |
| 総合ランキング | 84 | 69.0 | AA |
| マルチモーダルランキング | 17 | 64.0 | LS |
| 科学 | 137 | 68.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Chat
Tau2 Airline
67.0%自己申告
Communication
Tau2 Telecom
95.6%自己申告
Tau2 Retail
77.9%自己申告
Math
AIME 2025
94.0%自己申告
LiveBench
72.0%
FrontierMath
26.7%自己申告
Multimodal
MMMU
85.4%自己申告
Reasoning
BrowseComp Long Context 128k
90.0%自己申告
GPQANYU + Cohere + Anthropic (2023)
88.1%自己申告
SWE-Bench Verified
76.3%自己申告
AA評価指数
(Artificial Analysis)Math Index(Artificial Analysis)94.0
Aime 25(MAA (Mathematical Association of America))94.0
Gpqa(NYU + Cohere + Anthropic (2023))87.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))87.0
Livecodebench(UC Berkeley + MIT + Cornell (2024))86.8
Tau2(Sierra + U Toronto + Vector Institute (2025))81.9
Lcr(Artificial Analysis)80.0
Ifbench(Google Research (2023))72.9
Terminalbench V2 152.4
Coding Index(Artificial Analysis)49.4
Terminalbench Hard(Stanford × Laude Institute (2026))45.5
Hle(Center for AI Safety + Scale AI (2025))28.5
Intelligence Index(Artificial Analysis)24.7
Tau Banking15.9
LLM Statsカテゴリスコア
(LLM Stats (zeroeval))Multimodal90
Physics90
Search90
Healthcare90
Biology90
Chemistry90
Vision90
Reasoning80
Frontend Development80
General80
Code80
Communication80
Tool Calling80
Chat70
Math60
価格設定
入力価格$1.25 / 1Mトークン
出力価格$10 / 1Mトークン
混合価格(3:1)$3.438 / 1Mトークン
キャッシュ読み取り価格$0.125 / 1Mトークン
速度
トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s
プロバイダー価格ランキング
プロバイダー価格ランキング
2 プロバイダー
最安: OpenAI最高: Neon
プロバイダー入力出力
1OpenAI最安
$0
$0.00001
2Neon
$1.25
$10
このモデルの異なるAPIプロバイダー間の価格を比較。