メインコンテンツへスキップ

Qwen3.7 Max

AlibabaQwenProprietary

説明

Qwen3.7 Max is Alibaba Cloud Qwen Team's proprietary flagship model for agent-driven workflows. It is designed for coding agents, office automation, MCP and multi-agent orchestration, and long-horizon autonomous execution, with a 1 million token context window and up to 65,536 output tokens. Qwen reports strong agentic coding results including 69.7 on Terminal-Bench 2.0-Terminus, 80.4 on SWE-bench Verified, 60.6 on SWE-Pro, and 78.3 on SWE-Multilingual, alongside 92.4 on GPQA Diamond and 97.1 on HMMT 2026 Feb.

リリース日
2026-05-19
パラメータ
コンテキスト長
1.0M
モダリティ
text

能力レーダー

45
general
63
coding
92
reasoning
67
science
70
agents
0
multimodal

ランキング

ドメイン#順位スコアソース
エージェント能力5
66.0
LS
コーディングランキング55
84.0
AA
総合ランキング31
85.0
AA
数学的推論30
85.0
LB
推論26
83.0
LB
科学48
83.0
AA

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Agents

QwenSVG1608.00 / 2000自己申告
Kernel Bench L396.0%自己申告
SpreadSheetBench-v187.0%自己申告
CoWorkBench67.2%自己申告
MCP-Mark60.8%自己申告
QwenWorldBench57.3%自己申告
Finance Agent v248.4%
VITA-Bench47.9%自己申告

Chat

IFEvalGoogle Research (2023)94.3%自己申告

Code

QwenWebBench1568.00 / 2000自己申告
Claw-Eval65.2%自己申告
ZClawBench64.3%自己申告
SkillsBench59.2%自己申告
NL2Repo47.2%自己申告

General

MAXIFE89.2%自己申告
Include86.2%自己申告
NOVA-6359.0%自己申告

Instruction Following

IFBench79.1%自己申告

Language

MMLU-Redux95.0%自己申告
MMMLU90.3%自己申告
MMLU-Pro89.6%自己申告
MMLU-ProX87.0%自己申告
WMT24++85.8%自己申告

Long Context

MRCR 128K (8-needle)90.4%自己申告

Math

HMMT Feb 2697.1%自己申告
IMO-AnswerBench90.0%自己申告
PolyMATH86.5%自己申告
LiveBench74.3%
MathArena Apex44.5%自己申告

Reasoning

GPQANYU + Cohere + Anthropic (2023)92.4%自己申告
LiveCodeBench v691.6%自己申告
Global PIQA91.4%自己申告
SWE-Bench Verified80.4%自己申告
SWE-bench Multilingual78.3%自己申告
MCP Atlas76.4%自己申告
SuperGPQA73.6%自己申告
Terminal-Bench 2.0Stanford × Laude Institute (2026)69.7%自己申告
SWE-Bench ProPrinceton NLP (2024)60.6%自己申告
SciCode53.5%自己申告
Humanity's Last Exam41.4%自己申告
CritPT11.4%自己申告

Tool Calling

BFCL-V475.0%自己申告

AA評価指数

(Artificial Analysis)
Tau2(Sierra + U Toronto + Vector Institute (2025))
94.7
Gpqa(NYU + Cohere + Anthropic (2023))
92.3
Ifbench(Google Research (2023))
80.5
Lcr(Artificial Analysis)
74.7
Terminalbench V2 1
74.5
Coding Index(Artificial Analysis)
66.0
Terminalbench Hard(Stanford × Laude Institute (2026))
50.8
Scicode(UIUC + Argonne National Lab (2024))
48.8
Intelligence Index(Artificial Analysis)
46.7
Hle(Center for AI Safety + Scale AI (2025))
40.5
Tau Banking
11.8

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Chat
90
Language
90
Multimodal
90
Spatial Reasoning
90
Structured Output
90
Instruction Following
90
Legal
80
Physics
80
Productivity
80
Frontend Development
80
Healthcare
80
Math
70
Reasoning
70
Finance
70
General
70
Biology
70
Chemistry
70
Code
70
Economics
70
Tool Calling
70
Agents
60
Vision
60

価格設定

入力価格$2.5 / 1Mトークン
出力価格$7.5 / 1Mトークン
混合価格(3:1)$3.75 / 1Mトークン
キャッシュ読み取り価格$0.5 / 1Mトークン
キャッシュ書き込み価格$3.125 / 1Mトークン

速度

トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s

プロバイダー価格ランキング

プロバイダー価格ランキング

24 プロバイダー

最安: Novita最高: Modelis
プロバイダー入力出力
1Novita最安
$0
$0
2Together
$0
$0.00001
3Merge Gateway
$0.825
$2.4755
4NovitaAI
$1.25
$3.75
5Kilo Gateway
$1.25
$3.75
6DevPass (LLM Gateway)
$1.25
$3.75
7OrcaRouter
$1.25
$3.75
8Pioneer
$1.25
$3.75
9OpenRouter
$1.475
$4.425
10CrossModel
$1.504
$4.504
11AIHubMix
$1.69
$5.07
12Alibabaプライマリ
$2.5
$7.5
13NanoGPT
$2.5
$7.5
14Abacus
$2.5
$7.5
15OpenCode Go
$2.5
$7.5
16Alibaba (China)
$2.5
$7.5
17ZenMux
$2.5
$7.5
18Alibaba Coding Plan
$2.5
$7.5
19Requesty
$2.5
$7.5
20Alibaba Coding Plan (China)
$2.5
$7.5
21EmpirioLabs AI
$2.5
$7.5
22Charm Hyper
$2.5
$7.5
23Impossibl
$2.5
$7.5
24Modelis
$3
$9

このモデルの異なるAPIプロバイダー間の価格を比較。

外部リンク