Qwen3 VL 30B A3B Instruct
説明
Qwen3-VL is a large multimodal model that unifies vision, language, and reasoning to achieve human-level perception and cognition across text, images, and video. Built on a 235B-parameter architecture, it integrates early joint training of visual and textual modalities for strong language grounding. The model supports up to a 1 million-token context window and excels at visual understanding, spatial reasoning, long video comprehension, and tool-based interaction. It can generate code from images, perform precise 2D/3D object grounding, and operate digital interfaces like a visual agent. The “Instruct” version rivals Gemini 2.5 Pro in perception benchmarks, while the “Thinking” version leads in multimodal reasoning and STEM tasks. With multilingual OCR, creative writing, and fine-grained scene interpretation, Qwen3-VL establishes a new open-source frontier for integrated vision-language intelligence.
能力レーダー
ランキング
| ドメイン | #順位 | スコア | ソース |
|---|---|---|---|
| コーディングランキング | 427 | 31.0 | AA |
| 総合ランキング | 464 | 30.0 | AA |
| マルチモーダルランキング | 120 | 38.0 | LS |
| 科学 | 384 | 39.0 | AA |
ベンチマークスコア (LLM Stats)
(LLM Stats (zeroeval))Chat
Creativity
Factuality
General
Language
Math
Multimodal
Reasoning
Video
Vision
Writing
AA評価指数
(Artificial Analysis)LLM Statsカテゴリスコア
(LLM Stats (zeroeval))価格設定
速度
プロバイダー価格ランキング
プロバイダー価格ランキング
6 プロバイダー
このモデルの異なるAPIプロバイダー間の価格を比較。