Skip to main content

o1

OpenAIOpenAI o-seriesProprietary

Description

A research preview model focused on mathematical and logical reasoning capabilities, demonstrating improved performance on tasks requiring step-by-step reasoning, mathematical problem-solving, and code generation. The model shows enhanced capabilities in formal reasoning while maintaining strong general capabilities.

Release Date
2024-12-05
Parameters
—
Context Length
200K
Modalities
image, pdf, text

Capability Radar

35
general
51
coding
80
reasoning
54
science
60
agents
70
multimodal

Rankings

Domain#RankScoreSource
Code Ranking276
53.0
AA
General Ranking187
56.0
AA
Multimodal Ranking59
56.0
LS
Science330
43.0
AA

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Biology

GPQA Biology69.2%SR

Chat

TAU-bench Retail70.8%SR

Factuality

SimpleQA47.0%SR

General

MMLU91.8%SR

Language

MMMLU87.7%SR

Math

GSM8k97.1%SR
MATH96.4%SR
MGSM89.3%SR
AIME 202474.3%SR
MathVista71.8%SR
LiveBench67.0%SR
FrontierMath5.5%SR

Multimodal

MMMU77.6%SR

Reasoning

GPQA Physics92.8%SR
HumanEvalOpenAI (2021)88.1%SR
GPQANYU + Cohere + Anthropic (2023)78.0%SR
GPQA Chemistry64.7%SR
TAU-bench Airline50.0%SR
SWE-Bench Verified41.0%SR

AA Evaluation Indices

(Artificial Analysis)
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
97.0
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
84.1
Gpqa(NYU + Cohere + Anthropic (2023))
74.7
Aime(MAA (Mathematical Association of America))
72.3
Ifbench(Google Research (2023))
70.3
Livecodebench(UC Berkeley + MIT + Cornell (2024))
67.9
Lcr(Artificial Analysis)
65.0
Tau2(Sierra + U Toronto + Vector Institute (2025))
62.6
Coding Index(Artificial Analysis)
39.7
Intelligence Index(Artificial Analysis)
15.2
Terminalbench Hard(Stanford × Laude Institute (2026))
12.9
Hle(Center for AI Safety + Scale AI (2025))
7.0

LLM Stats Category Scores

(LLM Stats (zeroeval))
Language
90
Legal
90
Finance
90
Math
80
Physics
80
Healthcare
80
Biology
80
Chemistry
80
Chat
70
Multimodal
70
Reasoning
70
General
70
Vision
70
Code
60
Communication
60
Tool Calling
60
Factuality
50
Frontend Development
40

Pricing

Input Price$15 / 1M tokens
Output Price$60 / 1M tokens
Blended Price (3:1)$26.25 / 1M tokens
Cache Read Price$7.5 / 1M tokens

Speed

Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s

Provider Price Ranking

Provider Price Ranking

14 providers

Cheapest: PoeMost Expensive: LLM Gateway
ProviderInputOutput
1PoeCheapest
$14
$54
2OpenAIPRIMARY
$15
$60
3NanoGPT
$15
$60
4OpenRouter
$15
$60
5Kilo Gateway
$15
$60
6Helicone
$15
$60
7Azure Cognitive Services
$15
$60
8Vercel AI Gateway
$15
$60
9DevPass (LLM Gateway)
$15
$60
10Azure
$15
$60
11Merge Gateway
$15
$60
12Impossibl
$15
$60
13Eden AI
$15
$60
14LLM Gateway
$15
$60

Compare pricing across different API providers for this model.

External Sources