Qwen3.8 Max
AlibabaQwenOpen WeightQwen3.8-Max License · Commercial OK
Description
Qwen3.8 Max, released as the Qwen3.8-2.4T-A95B open-weight checkpoint, is Qwen's flagship mixture-of-experts model for coding, research, professional workflows, and long-horizon agents. It has 2.4 trillion total parameters with 95 billion active parameters, a native 262,144-token context window extensible to about 1 million tokens, and always-on thinking. QwenCloud serves the same model in the Qwen3.8 Max API family with text and image input.
Release Date
2026-08-03
Parameters
2.4T
Context Length
1.0M
Modalities
image, pdf, text, video
Capability Radar
55
general
69
coding
93
reasoning
69
science
60
agents
90
multimodal
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Agentic Capability | 4 | 79.0 | LS |
| Code Ranking | 26 | 91.0 | AA |
| General Ranking | 8 | 92.0 | AA |
| Math Reasoning | 17 | 91.0 | LB |
| Multimodal Ranking | 10 | 65.0 | LS |
| Reasoning | 13 | 88.0 | LB |
| Science | 26 | 87.0 | AA |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Agents
QwenReactBench
1724.00 / 2000SR
QwenSVG
1713.00 / 2000SR
PaperBench
93.0%SR
Terminal-Bench 2.1
86.6%SR
OSWorld-Verified
86.1%SR
AndroidWorld
85.3%SR
WideSearch
81.9%SR
QwenSWEBench
80.7%SR
MobileWorld
77.8%SR
AndroidBench
75.1%SR
CoWorkBench
74.8%SR
FrontierSWE
73.5%SR
Toolathlon
72.5%SR
SkillsBench
70.2%SR
Workspace Bench
67.7%SR
SWE-Bench ProPrinceton NLP (2024)
67.7%SR
QwenQoderBench
58.4%SR
DeepSWE 1.1
56.6%SR
NL2Repo
55.9%SR
Job Bench
53.4%SR
OneMillion Bench
52.5%SR
Agents' Last Exam
52.4%SR
MLS-Bench Lite
41.0%SR
AutomationBench
27.3%SR
Biology
GPQANYU + Cohere + Anthropic (2023)
92.6%SR
Code
Vision2Web
69.0%SR
Finance
PRBench-Finance
58.3%SR
General
MRCR v2 (8-needle)
92.9%SR
IFBench
82.8%SR
MMMU-Pro
82.3%SR
LongBench v2
66.3%SR
Grounding
ScreenSpot Pro
84.5%SR
Healthcare
HealthBench
60.2%SR
Knowledge
PLawBench
73.2%SR
PRBench-Legal
57.6%SR
Long Context
LVBench
81.8%SR
Math
Humanity's Last Exam (with tools, text-only)
56.2%SR
Humanity's Last Exam
43.6%SR
Multimodal
VideoMME w sub.
90.4%SR
PerceptionBench
63.5%SR
Reasoning
ERQA
77.8%SR
Spatial Reasoning
RealWorldQA
88.0%SR
AA Evaluation Indices
(Artificial Analysis)Coding Index(Artificial Analysis)71.8
Intelligence Index(Artificial Analysis)58.1
Gpqa(NYU + Cohere + Anthropic (2023))0.9
Terminalbench V2 10.8
Lcr(Artificial Analysis)0.7
Scicode(UIUC + Argonne National Lab (2024))0.5
Tau Banking0.5
Hle(Center for AI Safety + Scale AI (2025))0.4
LLM Stats Category Scores
(LLM Stats (zeroeval))Physics90
Biology90
Chemistry90
Video90
Multimodal80
Search80
Spatial Reasoning80
Instruction Following80
Grounding80
Vision80
Long Context70
Reasoning70
Structured Output70
General70
Agents70
Code70
Productivity60
Healthcare60
Tool Calling60
Math50
Pricing
Input Price$2 / 1M tokens
Output Price$6 / 1M tokens
Blended Price (3:1)$3 / 1M tokens
Cache Read Price$0.25 / 1M tokens
Cache Write Price$2.5 / 1M tokens
Speed
Tokens/sec45.2
Time to First Token1.95s
Time to Answer46.22s
Provider Price Ranking
Provider Price Ranking
13 providers
Cheapest: AIHubMixMost Expensive: Charm Hyper
ProviderInputOutput
1AIHubMixCheapest
$1.69
$5.07
2Alibaba (China)
$1.77744
$5.33231
3LLM Gateway
$1.815
$5.4461
4CrossModel
$1.88
$5.63
5AlibabaPRIMARY
$2
$6
6NanoGPT
$2
$6
7OpenRouter
$2
$6
8OpenCode Go
$2
$6
9Kilo Gateway
$2
$6
10DigitalOcean
$2
$6
11Merge Gateway
$2
$6
12EmpirioLabs AI
$2
$6
13Charm Hyper
$2
$6
Compare pricing across different API providers for this model.