Nova 2 Omni
AmazonAmazonProprietary
Description
Amazon Nova 2 Omni is Amazon's first unified multimodal reasoning model that processes text, documents, images, video, and audio inputs and generates both text and images from a single model, eliminating multi-model coordination complexity. It delivers strong multimodal perception, core reasoning, agentic tool use, and high-quality image generation and editing, with configurable extended thinking. It supports a 1M token context window, 200+ languages for text, and 10 languages for speech input.
Release Date
2025-12-02
Parameters
—
Context Length
—
Modalities
—
Capability Radar
70
general
0
coding
90
reasoning
68
scienceest.
70
agents
80
multimodal
Science is estimated from LLM Stats science scores or reasoning when no dedicated science benchmarks are available.
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Agentic Capability | 51 | 47.0 | LS |
| Multimodal Ranking | 94 | 34.0 | LS |
Benchmark Scores (LLM Stats)
(LLM Stats (zeroeval))Audio
MMAU
75.3%SR
CoVoST2
40.7%SR
Chat
Multi-Challenge
75.5%SR
Tau2 Airline
68.8%SR
Communication
Tau2 Telecom
80.0%SR
Tau2 Retail
78.3%SR
Instruction Following
IFBench
68.7%SR
Language
MMLU-Pro
80.7%SR
Math
AIME 2025
92.1%SR
Multimodal
RefCOCOg
86.3%SR
Video-MME
77.9%SR
QVHighlights
76.7%SR
MAVERIX
66.6%SR
RealKIE-FCC
59.8%SR
Tool Calling
BFCL-V4
58.3%SR
Vision
ScreenSpot
85.4%SR
MMMU-Pro
61.4%SR
OCRBench_V2
58.2%SR
AA Evaluation Indices
(Artificial Analysis)No AA evaluation data available
LLM Stats Category Scores
(LLM Stats (zeroeval))Math90
Spatial Reasoning90
Grounding90
Legal80
Reasoning80
Finance80
Healthcare80
Communication80
Video80
Chat70
Multimodal70
Instruction Following70
General70
Tool Calling70
Vision70
Language60
Image To Text60
Document Understanding60
Agents60
Audio60
Speech To Text40
Pricing
No pricing data available
Speed
No speed data available
Provider Price Ranking
No provider data available