Qwen3.8 Max (0902)
AlibabaQwenOpen WeightQwen3.8-Max License · Commercial OK
Description
Qwen3.8 Max, released as the Qwen3.8-2.4T-A95B open-weight checkpoint, is Qwen's flagship mixture-of-experts model for coding, research, professional workflows, and long-horizon agents. It has 2.4 trillion total parameters with 95 billion active parameters, a native 262,144-token context window extensible to about 1 million tokens, and always-on thinking. QwenCloud serves the same model in the Qwen3.8 Max API family with text and image input.
Release Date
2026-09-02
Parameters
2.4T
Context Length
1.0M
Modalities
image, pdf, text, video
Capability Radar
45
general
72
coding
93
reasoning
69
science
60
agents
90
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Agentic Capability | 18 | 61.0 | LS |
| Code Ranking | 42 | 92.0 | AA |
| General Ranking | 35 | 78.0 | AA |
| Math Reasoning | 31 | 91.0 | LB |
| Multimodal Ranking | 1 | 84.0 | LS |
| Reasoning | 24 | 88.0 | LB |
| Science | 61 | 80.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Agents
QwenSVG
1713.00 / 2000SR
AndroidWorld
85.3%SR
MobileWorld
77.8%SR
AndroidBench
75.1%SR
CoWorkBench
74.8%SR
Toolathlon
72.5%SR
Workspace Bench
67.7%SR
Job Bench
53.4%SR
OneMillion Bench
52.5%SR
Agents' Last Exam
52.4%SR
MLS-Bench Lite
41.0%SR
AutomationBench
27.3%SR
Code
QwenReactBench
1724.00 / 2000SR
PaperBench
93.0%SR
QwenSWEBench
80.7%SR
FrontierSWE
73.5%SR
SkillsBench
70.2%SR
Vision2Web
69.0%SR
QwenQoderBench
58.4%SR
DeepSWE 1.1
56.6%SR
NL2Repo
55.9%SR
Healthcare
HealthBench
60.2%SR
Instruction Following
IFBench
82.8%SR
Knowledge
PLawBench
73.2%SR
PRBench-Finance
58.3%SR
PRBench-Legal
57.6%SR
Long Context
MRCR v2 (8-needle)
92.9%SR
LongBench v2
66.3%SR
Multimodal
OSWorld-Verified
86.1%SR
Reasoning
GPQANYU + Cohere + Anthropic (2023)
92.6%SR
Terminal-Bench 2.1
86.6%SR
SWE-Bench ProPrinceton NLP (2024)
67.7%SR
Humanity's Last Exam (with tools, text-only)
56.2%SR
Humanity's Last Exam
43.6%SR
Search
WideSearch
81.9%SR
Vision
VideoMME w sub.
90.4%SR
RealWorldQA
88.0%SR
ScreenSpot Pro
84.5%SR
MMMU-Pro
82.3%SR
LVBench
81.8%SR
ERQA
77.8%SR
PerceptionBench
63.5%SR
AA Evaluation Indices
(Artificial Analysis)Gpqa(NYU + Cohere + Anthropic (2023))92.8
Terminalbench V2 188.8
Lcr(Artificial Analysis)80.3
Coding Index(Artificial Analysis)76.2
Scicode(UIUC + Argonne National Lab (2024))52.1
Tau Banking47.8
Intelligence Index(Artificial Analysis)45.4
Hle(Center for AI Safety + Scale AI (2025))43.1
Terminalbench V4 038.9
LLM Stats Category Scores
(LLM Stats (zeroeval))Physics90
Biology90
Chemistry90
Video90
Instruction Following80
Multimodal80
Search80
Spatial Reasoning80
Grounding80
Vision80
Long Context70
Reasoning70
Structured Output70
General70
Agents70
Code70
Productivity60
Healthcare60
Tool Calling60
Math50
Pricing
Input Price$2 / 1M tokens
Output Price$6 / 1M tokens
Blended Price (3:1)$3 / 1M tokens
Cache Read Price$0.25 / 1M tokens
Cache Write Price$2.5 / 1M tokens
Speed
Tokens/sec40.0
Time to First Token1.84s
Time to Answer51.88s
Provider Price Ranking
Provider Price Ranking
6 providers
Cheapest: DeepInfraMost Expensive: EmpirioLabs AI
ProviderInputOutput
1DeepInfraCheapest
$0
$0
2Novita
$0
$0.00001
3Fireworks
$0
$0.00001
4Together
$0
$0.00001
5AlibabaPRIMARY
$2
$6
6EmpirioLabs AI
$2
$6
Compare pricing across different API providers for this model.