GPT-5 (high)
OpenAIGPTProprietary
Description
GPT-5 is a flagship model from OpenAI designed for coding, reasoning, and agentic tasks across domains. It is optimized for coding and agentic tasks with higher reasoning capabilities and medium speed.
Release Date
2025-08-07
Parameters
—
Context Length
400K
Modalities
image, text
Capability Radar
50
general
55
coding
95
reasoning
59
science
80
agents
90
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Agentic Capability | 75 | 45.0 | LS |
| Code Ranking | 134 | 66.0 | AA |
| General Ranking | 68 | 76.0 | AA |
| Multimodal Ranking | 43 | 49.0 | LS |
| Science | 99 | 70.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Agents
BrowseCompOpenAI (2025)
54.9%SR
Biology
GPQANYU + Cohere + Anthropic (2023)
87.3%SR
Code
SWE-Lancer (IC-Diamond subset)
100.0%SR
HumanEvalOpenAI (2021)
93.4%SR
Aider-Polyglot
88.0%SR
SWE-Bench Verified
74.9%SR
Communication
Tau2 Telecom
96.7%SR
Tau2 Retail
81.1%SR
Multi-Challenge
69.6%SR
Tau2 Airline
62.6%SR
Finance
MMLU
92.5%SR
General
MMMU
84.2%SR
MMMU-Pro
78.4%SR
Internal API instruction following (hard)
64.0%SR
LongFact Objects
0.8%SR
LongFact Concepts
0.7%SR
Healthcare
VideoMMMU
84.6%SR
HealthBench Hard
1.6%SR
Language
COLLIE
99.0%SR
Long Context
OpenAI-MRCR: 2 needle 128k
95.2%SR
OpenAI-MRCR: 2 needle 256k
86.8%SR
Math
AIME 2025
94.6%SR
HMMT 2025
93.3%SR
MATH
84.7%SR
FrontierMath
26.3%SR
Humanity's Last Exam
24.8%SR
Multimodal
VideoMME w sub.
86.7%SR
CharXiv-R
81.1%SR
Reasoning
BrowseComp Long Context 128k
90.0%SR
BrowseComp Long Context 256k
88.8%SR
Graphwalks BFS <128k
78.3%SR
Graphwalks parents <128k
73.3%SR
ERQA
65.7%SR
FActScoreMin et al. (NYU/UW, 2023)
1.0%SR
AA Evaluation Indices
(Artificial Analysis)Math Index(Artificial Analysis)94.3
Coding Index(Artificial Analysis)37.8
Intelligence Index(Artificial Analysis)35.3
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))1.0
Aime(MAA (Mathematical Association of America))1.0
Aime 25(MAA (Mathematical Association of America))0.9
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))0.9
Gpqa(NYU + Cohere + Anthropic (2023))0.9
Tau2(Sierra + U Toronto + Vector Institute (2025))0.8
Livecodebench(UC Berkeley + MIT + Cornell (2024))0.8
Lcr(Artificial Analysis)0.8
Ifbench(Google Research (2023))0.7
Scicode(UIUC + Argonne National Lab (2024))0.4
Terminalbench V2 10.4
Terminalbench Hard(Stanford × Laude Institute (2026))0.3
Hle(Center for AI Safety + Scale AI (2025))0.3
Tau Banking0.2
LLM Stats Category Scores
(LLM Stats (zeroeval))Spatial Reasoning7
Vision4
Reasoning2
General1
Long Context100
Language100
Writing100
Legal90
Physics90
Finance90
Biology90
Chemistry90
Code90
Video90
Multimodal80
Communication80
Tool Calling80
Math70
Search70
Frontend Development70
Healthcare70
Structured Output60
Agents50
Pricing
Input Price$1.25 / 1M tokens
Output Price$10 / 1M tokens
Blended Price (3:1)$3.438 / 1M tokens
Cache Read Price$0.125 / 1M tokens
Speed
Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s
Provider Price Ranking
Provider Price Ranking
1 providers
ProviderInputOutput
1OpenAIPRIMARY
$1.25
$10
Compare pricing across different API providers for this model.