メインコンテンツへスキップ

o1

OpenAIOpenAI o-seriesProprietary

説明

A research preview model focused on mathematical and logical reasoning capabilities, demonstrating improved performance on tasks requiring step-by-step reasoning, mathematical problem-solving, and code generation. The model shows enhanced capabilities in formal reasoning while maintaining strong general capabilities.

リリース日
2024-12-05
パラメータ
コンテキスト長
200K
モダリティ
image, pdf, text

能力レーダー

39
general
49
coding
80
reasoning
48
science
60
agents
70
multimodal

ランキング

ドメイン#順位スコアソース
コーディングランキング200
54.0
AA
総合ランキング148
62.0
AA
科学250
49.0
AA

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Biology

GPQANYU + Cohere + Anthropic (2023)78.0%自己申告
GPQA Biology69.2%自己申告

Chemistry

GPQA Chemistry64.7%自己申告

Code

HumanEvalOpenAI (2021)88.1%自己申告
SWE-Bench Verified41.0%自己申告

Communication

TAU-bench Retail70.8%自己申告
TAU-bench Airline50.0%自己申告

Factuality

SimpleQA47.0%自己申告

Finance

MMLU91.8%自己申告

General

MMMLU87.7%自己申告
MMMU77.6%自己申告
LiveBench67.0%自己申告

Math

GSM8k97.1%自己申告
MATH96.4%自己申告
MGSM89.3%自己申告
AIME 202474.3%自己申告
MathVista71.8%自己申告
FrontierMath5.5%自己申告

Physics

GPQA Physics92.8%自己申告

AA評価指数

(Artificial Analysis)
Coding Index(Artificial Analysis)
39.7
Intelligence Index(Artificial Analysis)
23.9
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
1.0
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
0.8
Gpqa(NYU + Cohere + Anthropic (2023))
0.7
Aime(MAA (Mathematical Association of America))
0.7
Ifbench(Google Research (2023))
0.7
Livecodebench(UC Berkeley + MIT + Cornell (2024))
0.7
Lcr(Artificial Analysis)
0.6
Tau2(Sierra + U Toronto + Vector Institute (2025))
0.6
Scicode(UIUC + Argonne National Lab (2024))
0.4
Terminalbench Hard(Stanford × Laude Institute (2026))
0.1
Hle(Center for AI Safety + Scale AI (2025))
0.1

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Legal
90
Language
90
Finance
90
Math
80
Physics
80
Healthcare
80
Biology
80
Chemistry
80
Multimodal
70
Reasoning
70
General
70
Vision
70
Code
60
Communication
60
Tool Calling
60
Factuality
50
Frontend Development
40

価格設定

入力価格$15 / 1Mトークン
出力価格$60 / 1Mトークン
混合価格(3:1)$26.25 / 1Mトークン
キャッシュ読み取り価格$7.5 / 1Mトークン

速度

トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s

プロバイダー価格ランキング

プロバイダー価格ランキング

13 プロバイダー

最安: Poe最高: Impossibl
プロバイダー入力出力
1Poe最安
$14
$54
2OpenAIプライマリ
$15
$60
3NanoGPT
$15
$60
4OpenRouter
$15
$60
5Kilo Gateway
$15
$60
6Cloudflare AI Gateway
$15
$60
7Helicone
$15
$60
8Azure Cognitive Services
$15
$60
9Vercel AI Gateway
$15
$60
10LLM Gateway
$15
$60
11Azure
$15
$60
12Merge Gateway
$15
$60
13Impossibl
$15
$60

このモデルの異なるAPIプロバイダー間の価格を比較。

外部リンク