Qwen3 VL 235B A22B (Reasoning)
Описание
Qwen3-VL-235B-A22B-Thinking is the most powerful vision-language model in the Qwen series, featuring 236B parameters with MoE architecture for reasoning-enhanced multimodal understanding. Key capabilities include: Visual Agent (operates PC/mobile GUIs, recognizes elements, invokes tools), Visual Coding (generates Draw.io/HTML/CSS/JS from images/videos), Advanced Spatial Perception (2D grounding and 3D grounding for spatial reasoning and embodied AI), Long Context & Video Understanding (native 256K context expandable to 1M, handles hours-long video with second-level indexing), Enhanced Multimodal Reasoning (excels in STEM/Math with causal analysis), Upgraded Visual Recognition (celebrities, anime, products, landmarks, flora/fauna), and Expanded OCR (32 languages, robust in low light/blur/tilt). Architecture innovations include Interleaved-MRoPE for positional embeddings, DeepStack for multi-level ViT feature fusion, and Text-Timestamp Alignment for precise video temporal modeling.
Радар способностей
Рейтинги
| Домен | #Место | Оценка | Источник |
|---|---|---|---|
| Рейтинг кодинга | 274 | 53.0 | AA |
| Общий рейтинг | 231 | 50.0 | AA |
| Мультимодальный рейтинг | 81 | 51.0 | LS |
| Наука | 279 | 48.0 | AA |
Оценки бенчмарков (LLM Stats)
(LLM Stats (zeroeval))Chat
Creativity
Factuality
General
Language
Math
Multimodal
Reasoning
Video
Vision
Writing
Индексы оценки AA
(Artificial Analysis)Оценки категорий LLM Stats
(LLM Stats (zeroeval))Цены
Скорость
Рейтинг цен провайдеров
Рейтинг цен провайдеров
11 провайдеров
Сравнение цен разных API-провайдеров для этой модели.