Muse Spark
MetaProprietary
Description
Muse Spark is the first model in the Muse family developed by Meta Superintelligence Labs. It is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration. It features a Contemplating mode that orchestrates multiple agents reasoning in parallel. It demonstrates competitive performance in multimodal perception, reasoning, health, and agentic tasks, with Contemplating mode achieving 58% on Humanity's Last Exam and 38% on FrontierScience Research.
Release Date
2026-04-08
Parameters
—
Context Length
—
Modalities
—
Capability Radar
42
general
58
coding
88
reasoning
66
scienceest.
80
agents
70
multimodal
Science uses a reasoning proxy when dedicated science benchmarks are unavailable.
Rankings
| Domain | #Rank | Score | Source |
|---|---|---|---|
| Agentic Capability | 67 | 54.0 | LS |
| Code Ranking | 33 | 79.0 | AA |
| General Ranking | 19 | 81.0 | AA |
| Multimodal Ranking | 77 | 60.0 | LS |
| Reasoning | 92 | 50.0 | LS |
| Science | 15 | 84.0 | AA |
Benchmark Scores (LLM Stats)
Agents
GDPval-AA
1164.00 / 3000SR
DeepSearchQA
74.8%SR
Terminal-Bench 2.0
59.0%SR
SWE-Bench Pro
52.4%SR
Biology
GPQA
89.5%SR
Code
LiveCodeBench Pro
0.80 / 3000SR
SWE-Bench Verified
77.4%SR
Communication
Tau2 Telecom
91.5%SR
General
MMMU-Pro
80.4%SR
SimpleVQA
0.71 / 100SR
Grounding
ScreenSpot Pro
84.1%SR
Healthcare
MedXpertQA
78.4%SR
HealthBench Hard
42.8%SR
Math
Humanity's Last Exam
58.4%SR
Multimodal
CharXiv-R
86.4%SR
ZEROBench
0.33 / 100SR
Physics
IPhO 2025
82.6%SR
Reasoning
ERQA
64.7%SR
ARC-AGI v2
42.5%SR
FrontierScience Research
38.3%SR
AA Evaluation Indices
Coding Index58.6
Intelligence Index43.1
Tau20.9
Gpqa0.9
Ifbench0.8
Lcr0.7
Terminalbench V2 10.6
Scicode0.5
Terminalbench Hard0.5
Hle0.4
Tau Banking0.2
LLM Stats Category Scores
Legal100
Finance100
Agents100
General100
Reasoning78
Physics90
Biology90
Chemistry90
Communication90
Frontend Development80
Grounding80
Tool Calling80
Image To Text70
Multimodal70
Search70
Code70
Vision70
Math60
Spatial Reasoning60
Healthcare60
Pricing
Input PriceFree
Output PriceFree
Blended Price (3:1)Free
Speed
Tokens/sec0.0
Time to First Token0.00s
Time to Answer0.00s
Provider Price Ranking
No provider data available