Muse Spark
MetaProprietary
Описание
Muse Spark is the first model in the Muse family developed by Meta Superintelligence Labs. It is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration. It features a Contemplating mode that orchestrates multiple agents reasoning in parallel. It demonstrates competitive performance in multimodal perception, reasoning, health, and agentic tasks, with Contemplating mode achieving 58% on Humanity's Last Exam and 38% on FrontierScience Research.
Дата выхода
2026-04-08
Параметры
—
Длина контекста
—
Модальности
—
Радар способностей
44
general
58
coding
88
reasoning
66
science
80
agents
70
multimodal
Рейтинги
| Домен | #Место | Оценка | Источник |
|---|---|---|---|
| Агентные возможности | 128 | 34.0 | LS |
| Рейтинг кодинга | 67 | 78.0 | AA |
| Общий рейтинг | 42 | 81.0 | AA |
| Мультимодальный рейтинг | 8 | 65.0 | LS |
| Наука | 39 | 84.0 | AA |
Оценки бенчмарков (LLM Stats)
(LLM Stats (zeroeval))Agents
DeepSearchQA
74.8%Сам.
Terminal-Bench 2.0Stanford × Laude Institute (2026)
59.0%Сам.
SWE-Bench ProPrinceton NLP (2024)
52.4%Сам.
Biology
GPQANYU + Cohere + Anthropic (2023)
89.5%Сам.
Code
LiveCodeBench Pro
0.80 / 3000Сам.
SWE-Bench Verified
77.4%Сам.
Communication
Tau2 Telecom
91.5%Сам.
General
MMMU-Pro
80.4%Сам.
SimpleVQA
0.71 / 100Сам.
Grounding
ScreenSpot Pro
84.1%Сам.
Healthcare
MedXpertQA
78.4%Сам.
HealthBench Hard
42.8%Сам.
Math
Humanity's Last Exam
58.4%Сам.
Multimodal
CharXiv-R
86.4%Сам.
ZEROBench
0.33 / 100Сам.
Physics
IPhO 2025
82.6%Сам.
Reasoning
ERQA
64.7%Сам.
ARC-AGI v2
42.5%Сам.
FrontierScience Research
38.3%Сам.
Индексы оценки AA
(Artificial Analysis)Coding Index(Artificial Analysis)58.6
Intelligence Index(Artificial Analysis)44.3
Tau2(Sierra + U Toronto + Vector Institute (2025))0.9
Gpqa(NYU + Cohere + Anthropic (2023))0.9
Lcr(Artificial Analysis)0.8
Ifbench(Google Research (2023))0.8
Terminalbench V2 10.6
Scicode(UIUC + Argonne National Lab (2024))0.5
Terminalbench Hard(Stanford × Laude Institute (2026))0.5
Hle(Center for AI Safety + Scale AI (2025))0.4
Оценки категорий LLM Stats
(LLM Stats (zeroeval))Physics90
Biology90
Chemistry90
Communication90
Frontend Development80
Grounding80
Tool Calling80
Multimodal70
Reasoning70
Search70
Image To Text70
General70
Code70
Vision70
Math60
Spatial Reasoning60
Healthcare60
Agents60
Science40
Цены
Цена вводаБесплатно
Цена выводаБесплатно
Смешанная цена (3:1)Бесплатно
Скорость
Токенов/сек0.0
Задержка первого токена0.00s
Время до первого ответа0.00s
Рейтинг цен провайдеров
Нет данных провайдеров