GPT-4o (Aug '24)
OpenAIGPTProprietary
説明
GPT-4o ('o' for 'omni') is a multimodal AI model that accepts text, audio, image, and video inputs, and generates text, audio, and image outputs. It matches GPT-4 Turbo performance on text and code, with improvements in non-English languages, vision, and audio understanding.
リリース日
2024-08-06
パラメータ
—
コンテキスト長
128K
モダリティ
image, text
能力レーダー
7
general
32
coding
40
reasoning
37
science
50
agents
90
multimodal
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| コーディングランキング | 422 | 31.0 | AA |
| 総合ランキング | 571 | 20.0 | AA |
| マルチモーダルランキング | 115 | 39.0 | LS |
| 科学 | 501 | 26.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Chat
IFEvalGoogle Research (2023)
81.0%自己申告
Multi-IF
60.9%自己申告
TAU-bench Retail
60.3%自己申告
Tau2 Airline
45.5%自己申告
Multi-Challenge
40.3%自己申告
Communication
Tau2 Retail
63.4%自己申告
Tau2 Telecom
23.5%自己申告
Factuality
SimpleQA
38.2%自己申告
General
MMLU
85.7%自己申告
Aider-Polyglot
30.7%自己申告
Internal API instruction following (hard)
29.2%自己申告
Aider-Polyglot Edit
18.2%自己申告
Language
MMMLU
81.4%自己申告
MMLU-Pro
74.7%自己申告
COLLIE
61.0%自己申告
Long Context
ComplexFuncBench
66.5%自己申告
OpenAI-MRCR: 2 needle 128k
31.9%自己申告
Math
MathVista
61.4%自己申告
AIME 2024
13.1%自己申告
Multimodal
MMMU
72.2%自己申告
VideoMMMU
61.2%自己申告
Reasoning
ChartQAMasry et al. (2022)
85.7%自己申告
CharXiv-D
85.3%自己申告
GPQANYU + Cohere + Anthropic (2023)
70.1%自己申告
CharXiv-R
58.8%自己申告
TAU-bench Airline
42.8%自己申告
Graphwalks BFS <128k
41.7%自己申告
Graphwalks parents <128k
35.4%自己申告
SWE-Bench Verified
33.2%自己申告
SWE-Lancer
32.6%自己申告
SWE-Lancer (IC-Diamond subset)
12.4%自己申告
Humanity's Last Exam
5.3%自己申告
Vision
AI2D
94.2%自己申告
DocVQADocVQA (2020)
92.8%自己申告
EgoSchema
72.2%自己申告
ActivityNet
61.9%自己申告
MMMU-Pro
59.9%自己申告
ERQA
35.2%自己申告
AA評価指数
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))79.5
Gpqa(NYU + Cohere + Anthropic (2023))52.1
Lcr(Artificial Analysis)41.0
Ifbench(Google Research (2023))36.0
Livecodebench(UC Berkeley + MIT + Cornell (2024))31.7
Tau2(Sierra + U Toronto + Vector Institute (2025))28.9
Aime(MAA (Mathematical Association of America))11.7
Terminalbench Hard(Stanford × Laude Institute (2026))8.3
Intelligence Index(Artificial Analysis)7.7
Hle(Center for AI Safety + Scale AI (2025))2.3
LLM Statsカテゴリスコア
(LLM Stats (zeroeval))Image To Text90
Legal80
Finance80
Instruction Following70
Language70
Multimodal70
Physics70
Healthcare70
Biology70
Chemistry70
Vision70
Chat60
Long Context60
Structured Output60
Writing60
Math50
Reasoning50
General50
Communication50
Tool Calling50
Spatial Reasoning40
Factuality40
Frontend Development30
Code30
価格設定
入力価格$2.5 / 1Mトークン
出力価格$10 / 1Mトークン
混合価格(3:1)$4.375 / 1Mトークン
速度
トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s
プロバイダー価格ランキング
プロバイダー価格ランキング
2 プロバイダー
最安: OpenAI最高: Azure
プロバイダー入力出力
1OpenAI最安
$0
$0.00001
2Azure
$0
$0.00001
このモデルの異なるAPIプロバイダー間の価格を比較。