Skip to main content

Qwen2.5 Instruct 72B

AlibabaQwenOpen WeightQwen · Commercial OK

Description

Qwen2.5-72B-Instruct is an instruction-tuned 72 billion parameter language model, part of the Qwen2.5 series. It is designed to follow instructions, generate long texts (over 8K tokens), understand structured data (e.g., tables), and generate structured outputs, especially JSON. The model supports multilingual capabilities across over 29 languages.

Release Date
2024-09-19
Parameters
72.7B
Context Length
131K
Modalities
text

Capability Radar

26
general
28
coding
29
reasoning
35
science
29
agents
0
multimodal

Rankings

Domain#RankScoreSource
Code Ranking513
19.0
AA
General Ranking422
33.0
AA
Science517
25.0
AA

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Chat

MT-Bench0.94 / 100SR
IFEvalGoogle Research (2023)84.1%SR

General

AlignBench81.6%SR
Arena Hard81.2%SR
MultiPL-E75.1%SR

Language

MMLU-Redux86.8%SR
MMLU-Pro71.1%SR

Math

GSM8k95.8%SR
MATH83.1%SR
LiveBench52.3%SR

Reasoning

MBPP0.88 / 100SR
HumanEvalOpenAI (2021)86.6%SR
LiveCodeBench55.5%SR
GPQANYU + Cohere + Anthropic (2023)49.0%SR

AA Evaluation Indices

(Artificial Analysis)
Math 500(OpenAI (2024), subset of Hendrycks et al. MATH (2021))
85.8
Mmlu Pro(TIGER-Lab (Univ. of Waterloo, Toronto, CMU, 2024))
72.0
Gpqa(NYU + Cohere + Anthropic (2023))
49.1
Ifbench(Google Research (2023))
36.9
Tau2(Sierra + U Toronto + Vector Institute (2025))
34.5
Livecodebench(UC Berkeley + MIT + Cornell (2024))
27.6
Aime(MAA (Mathematical Association of America))
16.0
Aime 25(MAA (Mathematical Association of America))
14.0
Math Index(Artificial Analysis)
14.0
Intelligence Index(Artificial Analysis)
7.7
Terminalbench Hard(Stanford × Laude Institute (2026))
4.5
Hle(Center for AI Safety + Scale AI (2025))
3.6

LLM Stats Category Scores

(LLM Stats (zeroeval))
Chat
90
Roleplay
90
Communication
90
Creativity
90
Instruction Following
80
Language
80
Math
80
Reasoning
80
Structured Output
80
General
80
Writing
80
Legal
70
Finance
70
Healthcare
70
Code
70
Physics
50
Biology
50
Chemistry
50

Pricing

Input Price$0.475 / 1M tokens
Output Price$0.495 / 1M tokens
Blended Price (3:1)$0.48 / 1M tokens

Speed

Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s

Provider Price Ranking

Provider Price Ranking

3 providers

Cheapest: DeepInfraMost Expensive: Alibaba (China)
ProviderInputOutput
1DeepInfraCheapest
$0
$0
2AlibabaPRIMARY
$0.475
$0.495
3Alibaba (China)
$0.574
$1.721

Compare pricing across different API providers for this model.

External Sources