Skip to main content

MiMo-V2-Omni

XiaomiProprietary

Description

MiMo-V2-Omni is Xiaomi's omni foundation model uniting frontier multimodal understanding with strong agentic capability. It fuses dedicated image, video, and audio encoders into a single shared backbone, processing all modalities simultaneously. Natively supports structured tool calling, function execution, and UI grounding. Supports over 10 hours of continuous audio understanding and 256K token context window.

Release Date
2026-03-19
Parameters
—
Context Length
262K
Modalities
audio, image, pdf, text, video

Capability Radar

24
general
70
coding
83
reasoning
64
science
70
agents
85
multimodal

Rankings

Domain#RankScoreSource
Code Ranking188
69.0
AA
General Ranking175
57.0
AA
Science191
60.0
AA

Benchmark Scores (LLM Stats)

(LLM Stats (zeroeval))

Agents

MM-BrowserComp52.0%SR
OmniGAIA49.8%SR

Code

PinchBench81.2%SR
Claw-Eval54.8%SR

Reasoning

SWE-Bench Verified74.8%SR

AA Evaluation Indices

(Artificial Analysis)
Tau2(Sierra + U Toronto + Vector Institute (2025))
91.2
Gpqa(NYU + Cohere + Anthropic (2023))
82.8
Lcr(Artificial Analysis)
75.0
Ifbench(Google Research (2023))
53.5
Terminalbench Hard(Stanford × Laude Institute (2026))
34.8
Intelligence Index(Artificial Analysis)
23.9
Hle(Center for AI Safety + Scale AI (2025))
22.1

LLM Stats Category Scores

(LLM Stats (zeroeval))
Reasoning
70
Frontend Development
70
General
70
Agents
70
Code
70

Pricing

Input PriceFree
Output PriceFree
Blended Price (3:1)Free
Cache Read Price$0.0028 / 1M tokens

Speed

Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s

Provider Price Ranking

Provider Price Ranking

1 providers

ProviderInputOutput
1Xiaomi
$0.14
$0.28

Compare pricing across different API providers for this model.

External Sources