GPT-4.5 (Preview)
OpenAIGPTProprietary
描述
GPT-4.5 is OpenAI's most advanced model, offering improved reasoning, coding, and creative capabilities with faster performance and longer context handling than GPT-4. It features enhanced instruction following, reduced hallucinations, and better factual accuracy.
發布日期
2025-02-27
參數規模
—
上下文長度
—
支援模態
image, text
能力雷達圖
10
general
50
coding
80
reasoning
60
science估算
60
agents
70
multimodal
缺少專門科學評測時,Science 由 LLM Stats 科學得分或推理能力估算。
排行榜排名
基準測試分數 (LLM Stats)
(LLM Stats (zeroeval))Chat
IFEvalGoogle Research (2023)
88.2%自報
Multi-IF
70.8%自報
TAU-bench Retail
68.4%自報
Multi-Challenge
43.8%自報
Factuality
SimpleQA
62.5%自報
General
MMLU
90.8%自報
Internal API instruction following (hard)
54.0%自報
Aider-Polyglot Edit
44.9%自報
Language
MMMLU
85.1%自報
COLLIE
72.3%自報
Long Context
ComplexFuncBench
63.0%自報
OpenAI-MRCR: 2 needle 128k
38.5%自報
Math
GSM8k
97.0%自報
MathVista
72.3%自報
AIME 2024
36.7%自報
Multimodal
MMMU
75.2%自報
Reasoning
CharXiv-D
90.0%自報
HumanEvalOpenAI (2021)
88.0%自報
Graphwalks parents <128k
72.6%自報
Graphwalks BFS <128k
72.3%自報
GPQANYU + Cohere + Anthropic (2023)
69.5%自報
CharXiv-R
55.4%自報
TAU-bench Airline
50.0%自報
SWE-Bench Verified
38.0%自報
SWE-Lancer
37.3%自報
SWE-Lancer (IC-Diamond subset)
17.4%自報
AA 評測指數
(Artificial Analysis)Intelligence Index(Artificial Analysis)9.6
LLM Stats 分類評分
(LLM Stats (zeroeval))Legal90
Finance90
Instruction Following80
Language80
Math80
Healthcare80
Chat70
Multimodal70
Physics70
Spatial Reasoning70
Structured Output70
Biology70
Chemistry70
Vision70
Writing70
Reasoning60
Factuality60
General60
Communication60
Tool Calling60
Long Context50
Code50
Frontend Development40
定價
輸入價格免費
輸出價格免費
混合價格(3:1)免費
速度
Tokens/秒0.0
首Token延遲0.00s
首回答延遲0.00s
供應商價格排行
暫無提供商資料