GPT-4o (Aug '24)
OpenAIGPTProprietary
Description
GPT-4o ('o' for 'omni') is a multimodal AI model that accepts text, audio, image, and video inputs, and generates text, audio, and image outputs. It matches GPT-4 Turbo performance on text and code, with improvements in non-English languages, vision, and audio understanding.
Release Date
2024-08-06
Parameters
—
Context Length
128K
Modalities
image, pdf, text
Capability Radar
8
general
32
coding
40
reasoning
35
science
50
agents
90
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Code Ranking | 336 | 31.0 | AA |
| General Ranking | 478 | 23.0 | AA |
| Multimodal Ranking | 42 | 48.0 | LS |
| Science | 368 | 36.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Biology
GPQANYU + Cohere + Anthropic (2023)
70.1%SR
Code
SWE-Bench Verified
33.2%SR
SWE-Lancer
32.6%SR
Aider-Polyglot
30.7%SR
Aider-Polyglot Edit
18.2%SR
SWE-Lancer (IC-Diamond subset)
12.4%SR
Communication
Tau2 Retail
63.4%SR
Multi-IF
60.9%SR
TAU-bench Retail
60.3%SR
Tau2 Airline
45.5%SR
TAU-bench Airline
42.8%SR
Multi-Challenge
40.3%SR
Tau2 Telecom
23.5%SR
Factuality
SimpleQA
38.2%SR
Finance
MMLU
85.7%SR
MMLU-Pro
74.7%SR
General
MMMLU
81.4%SR
IFEvalGoogle Research (2023)
81.0%SR
MMMU
72.2%SR
MMMU-Pro
59.9%SR
Internal API instruction following (hard)
29.2%SR
Healthcare
VideoMMMU
61.2%SR
Image To Text
DocVQADocVQA (2020)
92.8%SR
Language
COLLIE
61.0%SR
Long Context
EgoSchema
72.2%SR
ComplexFuncBench
66.5%SR
OpenAI-MRCR: 2 needle 128k
31.9%SR
Math
MathVista
61.4%SR
AIME 2024
13.1%SR
Humanity's Last Exam
5.3%SR
Multimodal
AI2D
94.2%SR
ChartQAMasry et al. (2022)
85.7%SR
CharXiv-D
85.3%SR
CharXiv-R
58.8%SR
Reasoning
Graphwalks BFS <128k
41.7%SR
Graphwalks parents <128k
35.4%SR
ERQA
35.2%SR
Video
ActivityNet
61.9%SR
AA Evaluation Indices
(Artificial Analysis)Intelligence Index(Artificial Analysis)9.4
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))0.8
Gpqa(NYU + Cohere + Anthropic (2023))0.5
Lcr(Artificial Analysis)0.4
Ifbench(Google Research (2023))0.4
Scicode(UIUC + Argonne National Lab (2024))0.3
Livecodebench(UC Berkeley + MIT + Cornell (2024))0.3
Tau2(Sierra + U Toronto + Vector Institute (2025))0.3
Aime(MAA (Mathematical Association of America))0.1
Terminalbench Hard(Stanford × Laude Institute (2026))0.1
Hle(Center for AI Safety + Scale AI (2025))0.0
LLM Stats Category Scores
(LLM Stats (zeroeval))Image To Text90
Legal80
Finance80
Instruction Following70
Language70
Multimodal70
Physics70
Healthcare70
Biology70
Chemistry70
Vision70
Long Context60
Structured Output60
Writing60
Math50
Reasoning50
General50
Communication50
Tool Calling50
Spatial Reasoning40
Factuality40
Frontend Development30
Code30
Pricing
Input Price$2.5 / 1M tokens
Output Price$10 / 1M tokens
Blended Price (3:1)$4.375 / 1M tokens
Cache Read Price$1.25 / 1M tokens
Speed
Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s
Provider Price Ranking
Provider Price Ranking
2 providers
Cheapest: OpenAIMost Expensive: Azure
ProviderInputOutput
1OpenAICheapest
$0
$0.00001
2Azure
$0
$0.00001
Compare pricing across different API providers for this model.