メインコンテンツへスキップ

GPT-4o (Aug '24)

OpenAIGPTProprietary

説明

GPT-4o ('o' for 'omni') is a multimodal AI model that accepts text, audio, image, and video inputs, and generates text, audio, and image outputs. It matches GPT-4 Turbo performance on text and code, with improvements in non-English languages, vision, and audio understanding.

リリース日
2024-08-06
パラメータ
—
コンテキスト長
128K
モダリティ
image, text

能力レーダー

7
general
32
coding
40
reasoning
37
science
50
agents
90
multimodal

ランキング

ドメイン#順位スコアソース
コーディングランキング422
31.0
AA
総合ランキング571
20.0
AA
マルチモーダルランキング115
39.0
LS
科学501
26.0
AA

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Chat

IFEvalGoogle Research (2023)81.0%自己申告
Multi-IF60.9%自己申告
TAU-bench Retail60.3%自己申告
Tau2 Airline45.5%自己申告
Multi-Challenge40.3%自己申告

Communication

Tau2 Retail63.4%自己申告
Tau2 Telecom23.5%自己申告

Factuality

SimpleQA38.2%自己申告

General

MMLU85.7%自己申告
Aider-Polyglot30.7%自己申告
Internal API instruction following (hard)29.2%自己申告
Aider-Polyglot Edit18.2%自己申告

Language

MMMLU81.4%自己申告
MMLU-Pro74.7%自己申告
COLLIE61.0%自己申告

Long Context

ComplexFuncBench66.5%自己申告
OpenAI-MRCR: 2 needle 128k31.9%自己申告

Math

MathVista61.4%自己申告
AIME 202413.1%自己申告

Multimodal

MMMU72.2%自己申告
VideoMMMU61.2%自己申告

Reasoning

ChartQAMasry et al. (2022)85.7%自己申告
CharXiv-D85.3%自己申告
GPQANYU + Cohere + Anthropic (2023)70.1%自己申告
CharXiv-R58.8%自己申告
TAU-bench Airline42.8%自己申告
Graphwalks BFS <128k41.7%自己申告
Graphwalks parents <128k35.4%自己申告
SWE-Bench Verified33.2%自己申告
SWE-Lancer32.6%自己申告
SWE-Lancer (IC-Diamond subset)12.4%自己申告
Humanity's Last Exam5.3%自己申告

Vision

AI2D94.2%自己申告
DocVQADocVQA (2020)92.8%自己申告
EgoSchema72.2%自己申告
ActivityNet61.9%自己申告
MMMU-Pro59.9%自己申告
ERQA35.2%自己申告

AA評価指数

(Artificial Analysis)
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
79.5
Gpqa(NYU + Cohere + Anthropic (2023))
52.1
Lcr(Artificial Analysis)
41.0
Ifbench(Google Research (2023))
36.0
Livecodebench(UC Berkeley + MIT + Cornell (2024))
31.7
Tau2(Sierra + U Toronto + Vector Institute (2025))
28.9
Aime(MAA (Mathematical Association of America))
11.7
Terminalbench Hard(Stanford × Laude Institute (2026))
8.3
Intelligence Index(Artificial Analysis)
7.7
Hle(Center for AI Safety + Scale AI (2025))
2.3

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Image To Text
90
Legal
80
Finance
80
Instruction Following
70
Language
70
Multimodal
70
Physics
70
Healthcare
70
Biology
70
Chemistry
70
Vision
70
Chat
60
Long Context
60
Structured Output
60
Writing
60
Math
50
Reasoning
50
General
50
Communication
50
Tool Calling
50
Spatial Reasoning
40
Factuality
40
Frontend Development
30
Code
30

価格設定

入力価格$2.5 / 1Mトークン
出力価格$10 / 1Mトークン
混合価格(3:1)$4.375 / 1Mトークン

速度

トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s

プロバイダー価格ランキング

プロバイダー価格ランキング

2 プロバイダー

最安: OpenAI最高: Azure
プロバイダー入力出力
1OpenAI最安
$0
$0.00001
2Azure
$0
$0.00001

このモデルの異なるAPIプロバイダー間の価格を比較。

外部リンク