Qwen3 VL 235B A22B (Reasoning)
説明
Qwen3-VL-235B-A22B-Thinking is the most powerful vision-language model in the Qwen series, featuring 236B parameters with MoE architecture for reasoning-enhanced multimodal understanding. Key capabilities include: Visual Agent (operates PC/mobile GUIs, recognizes elements, invokes tools), Visual Coding (generates Draw.io/HTML/CSS/JS from images/videos), Advanced Spatial Perception (2D grounding and 3D grounding for spatial reasoning and embodied AI), Long Context & Video Understanding (native 256K context expandable to 1M, handles hours-long video with second-level indexing), Enhanced Multimodal Reasoning (excels in STEM/Math with causal analysis), Upgraded Visual Recognition (celebrities, anime, products, landmarks, flora/fauna), and Expanded OCR (32 languages, robust in low light/blur/tilt). Architecture innovations include Interleaved-MRoPE for positional embeddings, DeepStack for multi-level ViT feature fusion, and Text-Timestamp Alignment for precise video temporal modeling.
能力レーダー
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| コーディングランキング | 279 | 53.0 | AA |
| 総合ランキング | 235 | 50.0 | AA |
| マルチモーダルランキング | 82 | 51.0 | LS |
| 科学 | 284 | 48.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Chat
Creativity
Factuality
General
Language
Math
Multimodal
Reasoning
Video
Vision
Writing
AA評価指数
(Artificial Analysis)LLM Statsカテゴリスコア
(LLM Stats (zeroeval))価格設定
速度
プロバイダー価格ランキング
プロバイダー価格ランキング
11 プロバイダー
このモデルの異なるAPIプロバイダー間の価格を比較。