Перейти к основному содержанию

Nemotron 3 Ultra 550B A55B (Reasoning)

NVIDIAОткрытые весаOpenMDW License v1.1 · Коммерческое использование

Описание

Nemotron 3 Ultra is NVIDIA's frontier-scale open model with 550B total / 55B active parameters, built for agentic reasoning, long-context analysis, tool use, and high-stakes RAG. It uses a hybrid Latent Mixture-of-Experts (LatentMoE) architecture interleaving Mamba-2, MoE, and select Attention layers, with Multi-Token Prediction (MTP) for native speculative decoding, and is pre-trained on ~20T tokens with an NVFP4 recipe. Reasoning is configurable on/off (plus a medium-effort mode) via the chat template. It supports up to a 1M-token context and 10 languages (English, French, Spanish, Italian, German, Japanese, Hindi, Korean, Brazilian Portuguese, Chinese). Released with open weights, training data, and recipes under the OpenMDW-1.1 license.

Дата выхода
2026-06-04
Параметры
550.0B
Длина контекста
1.0M
Модальности
text

Радар способностей

36
general
48
coding
87
reasoning
59
science
40
agents
0
multimodal

Рейтинги

Домен#МестоОценкаИсточник
Агентные возможности93
26.0
LS
Рейтинг кодинга156
65.0
AA
Общий рейтинг85
74.0
AA
Наука121
67.0
AA

Оценки бенчмарков (LLM Stats)

(LLM Stats (zeroeval))

Agents

Finance Agent53.7%Сам.
Finance Agent v237.5%
TAU3-Bench22.6%Сам.

Chat

Multi-Challenge63.8%Сам.

Code

PinchBench90.0%Сам.

General

GDPval46.7%Сам.

Instruction Following

IFBench81.7%Сам.

Language

MMLU-Pro86.8%Сам.
WMT24++83.7%Сам.
MMLU-ProX83.0%Сам.

Long Context

RULER94.7%Сам.
LongBench v261.9%Сам.

Math

IMO-AnswerBench92.3%Сам.

Reasoning

LiveCodeBench v689.0%Сам.
GPQANYU + Cohere + Anthropic (2023)87.0%Сам.
Apex84.8%Сам.
SWE-Bench Verified70.7%Сам.
SWE-bench Multilingual67.7%Сам.
AA-LCR65.4%Сам.
Terminal-Bench 2.156.4%Сам.
ProfBench56.0%Сам.
SciCode44.6%Сам.
BrowseCompOpenAI (2025)44.4%Сам.
Humanity's Last Exam37.4%Сам.
CritPT3.1%Сам.

Science

OmniScience78.7%Сам.

Индексы оценки AA

(Artificial Analysis)
Gpqa(NYU + Cohere + Anthropic (2023))
86.7
Tau2(Sierra + U Toronto + Vector Institute (2025))
83.3
Ifbench(Google Research (2023))
81.4
Lcr(Artificial Analysis)
71.0
Terminalbench V2 1
53.9
Coding Index(Artificial Analysis)
49.3
Scicode(UIUC + Argonne National Lab (2024))
39.9
Intelligence Index(Artificial Analysis)
38.3
Terminalbench Hard(Stanford × Laude Institute (2026))
36.4
Hle(Center for AI Safety + Scale AI (2025))
28.4
Tau Banking
14.2

Оценки категорий LLM Stats

(LLM Stats (zeroeval))
Language
80
Science
80
Instruction Following
80
Healthcare
80
Legal
70
Long Context
70
Physics
70
Frontend Development
70
Biology
70
Chemistry
70
Code
70
Knowledge
70
Chat
60
Math
60
Reasoning
60
Structured Output
60
Finance
60
General
60
Communication
60
Agents
50
Search
40
Tool Calling
40
Vision
40

Цены

Цена ввода$0.675 / 1M токенов
Цена вывода$2.675 / 1M токенов
Смешанная цена (3:1)$1.175 / 1M токенов
Цена чтения кэша$0.15 / 1M токенов

Скорость

Токенов/сек101.7
Задержка первого токена2.57s
Время до первого ответа24.95s

Рейтинг цен провайдеров

Рейтинг цен провайдеров

7 провайдеров

Самый дешевый: NvidiaСамый дорогой: Venice AI
ПровайдерВводВывод
1NvidiaСамый дешевый
$0.5
$2.5
2NanoGPT
$0.5
$2.5
3Kilo Gateway
$0.5
$2.2
4Vercel AI Gateway
$0.6
$2.4
5Together AI
$0.6
$3.6
6OpenRouter
$0.625
$3.125
7Venice AI
$0.625
$3.125

Сравнение цен разных API-провайдеров для этой модели.

Внешние ссылки