Qwen3 Next 80B A3B (Reasoning)
説明
Qwen3-Next-80B-A3B-Thinking is the thinking variant of the Qwen3-Next series, featuring the same groundbreaking architecture as the instruct model. Leveraging GSPO, it addresses stability and efficiency challenges of hybrid attention + high-sparsity MoE in RL training. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared), and Multi-Token Prediction. With 80B total parameters and only 3B activated, it demonstrates outstanding performance on complex reasoning tasks — outperforming Qwen3-30B-A3B-Thinking-2507, Qwen3-32B-Thinking, and even the proprietary Gemini-2.5-Flash-Thinking across multiple benchmarks. Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)). Supports only thinking mode with automatic <think> tag inclusion, may generate longer thinking content.
能力レーダー
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| エージェント能力 | 46 | 53.0 | LS |
| コーディングランキング | 264 | 44.0 | AA |
| 総合ランキング | 233 | 51.0 | AA |
| 科学 | 207 | 54.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Agents
Biology
Chemistry
Code
Communication
Creativity
Finance
General
Math
Reasoning
AA評価指数
(Artificial Analysis)LLM Statsカテゴリスコア
(LLM Stats (zeroeval))価格設定
速度
プロバイダー価格ランキング
プロバイダー価格ランキング
10 プロバイダー
このモデルの異なるAPIプロバイダー間の価格を比較。