Skip to main content

Qwen3.5 397B A17B (Non-reasoning)

AlibabaQwenOpen WeightApache 2.0 · Commercial OK

Description

Qwen3.5-397B-A17B is Qwen's flagship Mixture-of-Experts model with 397 billion total parameters and 17 billion activated parameters. It delivers state-of-the-art performance across knowledge, reasoning, coding, mathematics, multilingual understanding, instruction following, long context, and agent tasks.

Release Date
2026-02-16
Parameters
397.0B
Context Length
262K
Modalities
audio, image, text, video

Capability Radar

21
general
70
coding
86
reasoning
66
science
60
agents
70
multimodal

Rankings

Domain#RankScoreSource
Agentic Capability17
62.0
LS
Code Ranking218
64.0
AA
General Ranking212
52.0
AA
Science196
60.0
AA

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Agents

t2-bench86.7%SR
VITA-Bench49.7%SR
MCP-Mark46.1%SR
Toolathlon38.3%SR
DeepPlanning34.3%SR

Chat

IFEvalGoogle Research (2023)92.6%SR
Multi-Challenge67.6%SR

Code

SecCodeBench68.3%SR

General

C-Eval93.0%SR
MAXIFE88.2%SR
Include85.6%SR
NOVA-6359.1%SR

Instruction Following

IFBench76.5%SR

Language

MMLU-Redux94.9%SR
MMMLU88.5%SR
MMLU-Pro87.8%SR
MMLU-ProX84.7%SR
WMT24++78.9%SR

Long Context

LongBench v263.2%SR

Math

HMMT 202594.8%SR
HMMT2592.7%SR
AIME 202691.3%SR
IMO-AnswerBench80.9%SR
PolyMATH73.3%SR

Reasoning

Global PIQA89.8%SR
GPQANYU + Cohere + Anthropic (2023)88.4%SR
LiveCodeBench v683.6%SR
SWE-Bench Verified76.4%SR
SuperGPQA70.4%SR
BrowseComp-zh70.3%SR
SWE-bench Multilingual69.3%SR
BrowseCompOpenAI (2025)69.0%SR
AA-LCR68.7%SR
Terminal-Bench 2.0Stanford × Laude Institute (2026)52.5%SR
Seal-046.9%SR
Humanity's Last Exam28.7%SR

Search

WideSearch74.0%SR

Tool Calling

BFCL-V472.9%SR

AA Evaluation Indices

(Artificial Analysis)
Gpqa(NYU + Cohere + Anthropic (2023))
86.1
Tau2(Sierra + U Toronto + Vector Institute (2025))
83.9
Lcr(Artificial Analysis)
64.3
Ifbench(Google Research (2023))
51.6
Terminalbench Hard(Stanford × Laude Institute (2026))
35.6
Intelligence Index(Artificial Analysis)
21.4
Hle(Center for AI Safety + Scale AI (2025))
19.8

LLM Stats Category Scores

(LLM Stats (zeroeval))
Language
90
Biology
90
Chat
80
Instruction Following
80
Legal
80
Math
80
Physics
80
Structured Output
80
Finance
80
Frontend Development
80
Healthcare
80
Chemistry
80
Long Context
70
Multimodal
70
Reasoning
70
Search
70
Spatial Reasoning
70
General
70
Code
70
Communication
70
Economics
70
Agents
60
Tool Calling
60
Vision
50

Pricing

Input Price$0.6 / 1M tokens
Output Price$3.6 / 1M tokens
Blended Price (3:1)$1.35 / 1M tokens

Speed

Tokens/sec85.4
Time to First Token1.59s
Time to Answer1.59s

Provider Price Ranking

Provider Price Ranking

2 providers

Cheapest: DeepInfraMost Expensive: Alibaba
ProviderInputOutput
1DeepInfraCheapest
$0
$0
2AlibabaPRIMARY
$0.6
$3.6

Compare pricing across different API providers for this model.

External Sources