GPT-5 (high)
OpenAIGPTProprietary
विवरण
GPT-5 is a flagship model from OpenAI designed for coding, reasoning, and agentic tasks across domains. It is optimized for coding and agentic tasks with higher reasoning capabilities and medium speed.
रिलीज़ तिथि
2025-08-07
पैरामीटर
—
संदर्भ लंबाई
400K
मोडैलिटीज़
image, text
क्षमता रडार
43
general
56
coding
95
reasoning
68
science
80
agents
90
multimodal
रैंकिंग
| डोमेन | #रैंक | स्कोर | स्रोत |
|---|---|---|---|
| कोडिंग रैंकिंग | 212 | 64.0 | AA |
| सामान्य रैंकिंग | 85 | 69.0 | AA |
| मल्टीमॉडल रैंकिंग | 20 | 63.0 | LS |
| विज्ञान | 150 | 66.0 | AA |
बेंचमार्क स्कोर (LLM Stats)
(LLM Stats (zeroeval))Chat
Multi-Challenge
69.6%स्वयं
Tau2 Airline
62.6%स्वयं
Communication
Tau2 Telecom
96.7%स्वयं
Tau2 Retail
81.1%स्वयं
General
MMLU
92.5%स्वयं
Aider-Polyglot
88.0%स्वयं
Internal API instruction following (hard)
64.0%स्वयं
LongFact Objects
0.8%स्वयं
LongFact Concepts
0.7%स्वयं
Healthcare
HealthBench Hard
1.6%स्वयं
Language
COLLIE
99.0%स्वयं
Long Context
OpenAI-MRCR: 2 needle 128k
95.2%स्वयं
OpenAI-MRCR: 2 needle 256k
86.8%स्वयं
Math
AIME 2025
94.6%स्वयं
HMMT 2025
93.3%स्वयं
MATH
84.7%स्वयं
FrontierMath
26.3%स्वयं
Multimodal
VideoMMMU
84.6%स्वयं
MMMU
84.2%स्वयं
Reasoning
SWE-Lancer (IC-Diamond subset)
100.0%स्वयं
HumanEvalOpenAI (2021)
93.4%स्वयं
BrowseComp Long Context 128k
90.0%स्वयं
BrowseComp Long Context 256k
88.8%स्वयं
GPQANYU + Cohere + Anthropic (2023)
87.3%स्वयं
CharXiv-R
81.1%स्वयं
Graphwalks BFS <128k
78.3%स्वयं
SWE-Bench Verified
74.9%स्वयं
Graphwalks parents <128k
73.3%स्वयं
BrowseCompOpenAI (2025)
54.9%स्वयं
Humanity's Last Exam
24.8%स्वयं
FActScoreMin et al. (NYU/UW, 2023)
1.0%स्वयं
Vision
VideoMME w sub.
86.7%स्वयं
MMMU-Pro
78.4%स्वयं
ERQA
65.7%स्वयं
AA मूल्यांकन सूचकांक
(Artificial Analysis)Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))99.4
Aime(MAA (Mathematical Association of America))95.7
Aime 25(MAA (Mathematical Association of America))94.3
Math Index(Artificial Analysis)94.3
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))87.1
Gpqa(NYU + Cohere + Anthropic (2023))85.4
Tau2(Sierra + U Toronto + Vector Institute (2025))84.8
Livecodebench(UC Berkeley + MIT + Cornell (2024))84.6
Lcr(Artificial Analysis)78.2
Ifbench(Google Research (2023))73.1
Coding Index(Artificial Analysis)37.8
Terminalbench V2 135.2
Terminalbench Hard(Stanford × Laude Institute (2026))32.6
Hle(Center for AI Safety + Scale AI (2025))28.5
Intelligence Index(Artificial Analysis)23.0
Tau Banking22.1
LLM Stats श्रेणी स्कोर
(LLM Stats (zeroeval))Spatial Reasoning7
Vision4
Reasoning2
General1
Language100
Long Context100
Writing100
Legal90
Physics90
Finance90
Biology90
Chemistry90
Code90
Video90
Multimodal80
Communication80
Tool Calling80
Chat70
Math70
Search70
Frontend Development70
Healthcare70
Structured Output60
Agents50
मूल्य निर्धारण
इनपुट मूल्य$1.25 / 1M टोकन
आउटपुट मूल्य$10 / 1M टोकन
मिश्रित मूल्य (3:1)$3.438 / 1M टोकन
कैश पठन मूल्य$0.125 / 1M टोकन
गति
टोकन/सेकंड0.0
पहले टोकन में देरी0.00s
पहले उत्तर में देरी0.00s
प्रदाता मूल्य रैंकिंग
प्रदाता मूल्य रैंकिंग
1 प्रदाता
प्रदाताइनपुटआउटपुट
1OpenAIप्राथमिक
$1.25
$10
इस मॉडल के लिए विभिन्न API प्रदाताओं के मूल्य निर्धारण की तुलना करें।