メインコンテンツへスキップ

Muse Spark

MetaProprietary

説明

Muse Spark is the first model in the Muse family developed by Meta Superintelligence Labs. It is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration. It features a Contemplating mode that orchestrates multiple agents reasoning in parallel. It demonstrates competitive performance in multimodal perception, reasoning, health, and agentic tasks, with Contemplating mode achieving 58% on Humanity's Last Exam and 38% on FrontierScience Research.

リリース日
2026-04-08
パラメータ
コンテキスト長
モダリティ

能力レーダー

44
general
58
coding
88
reasoning
66
science
80
agents
70
multimodal

ランキング

ベンチマークスコア (LLM Stats)

(LLM Stats (zeroeval))

Agents

DeepSearchQA74.8%自己申告
Terminal-Bench 2.0Stanford × Laude Institute (2026)59.0%自己申告
SWE-Bench ProPrinceton NLP (2024)52.4%自己申告

Biology

GPQANYU + Cohere + Anthropic (2023)89.5%自己申告

Code

LiveCodeBench Pro0.80 / 3000自己申告
SWE-Bench Verified77.4%自己申告

Communication

Tau2 Telecom91.5%自己申告

General

MMMU-Pro80.4%自己申告
SimpleVQA0.71 / 100自己申告

Grounding

ScreenSpot Pro84.1%自己申告

Healthcare

MedXpertQA78.4%自己申告
HealthBench Hard42.8%自己申告

Math

Humanity's Last Exam58.4%自己申告

Multimodal

CharXiv-R86.4%自己申告
ZEROBench0.33 / 100自己申告

Physics

IPhO 202582.6%自己申告

Reasoning

ERQA64.7%自己申告
ARC-AGI v242.5%自己申告
FrontierScience Research38.3%自己申告

AA評価指数

(Artificial Analysis)
Coding Index(Artificial Analysis)
58.6
Intelligence Index(Artificial Analysis)
44.3
Tau2(Sierra + U Toronto + Vector Institute (2025))
0.9
Gpqa(NYU + Cohere + Anthropic (2023))
0.9
Lcr(Artificial Analysis)
0.8
Ifbench(Google Research (2023))
0.8
Terminalbench V2 1
0.6
Scicode(UIUC + Argonne National Lab (2024))
0.5
Terminalbench Hard(Stanford × Laude Institute (2026))
0.5
Hle(Center for AI Safety + Scale AI (2025))
0.4

LLM Statsカテゴリスコア

(LLM Stats (zeroeval))
Physics
90
Biology
90
Chemistry
90
Communication
90
Frontend Development
80
Grounding
80
Tool Calling
80
Image To Text
70
Multimodal
70
Reasoning
70
Search
70
General
70
Code
70
Vision
70
Math
60
Spatial Reasoning
60
Healthcare
60
Agents
60
Science
40

価格設定

入力価格無料
出力価格無料
混合価格(3:1)無料

速度

トークン/秒0.0
初トークン遅延0.00s
初回答遅延0.00s

プロバイダー価格ランキング

プロバイダーデータがありません

外部リンク