메인 콘텐츠로 건너뛰기

Muse Spark

MetaProprietary

설명

Muse Spark is the first model in the Muse family developed by Meta Superintelligence Labs. It is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration. It features a Contemplating mode that orchestrates multiple agents reasoning in parallel. It demonstrates competitive performance in multimodal perception, reasoning, health, and agentic tasks, with Contemplating mode achieving 58% on Humanity's Last Exam and 38% on FrontierScience Research.

출시일
2026-04-08
파라미터
—
컨텍스트 길이
—
모달리티
—

능력 레이더

33
general
59
coding
88
reasoning
74
science
80
agents
70
multimodal

랭킹

도메인#순위점수소스
코딩 랭킹143
75.0
AA
종합 랭킹60
72.0
AA
멀티모달 랭킹12
66.0
LS
과학67
79.0
AA

벤치마크 점수 (LLM Stats)

(LLM Stats (zeroeval))

Communication

Tau2 Telecom91.5%자체 보고

Healthcare

MedXpertQA78.4%자체 보고
HealthBench Hard42.8%자체 보고

Reasoning

GPQANYU + Cohere + Anthropic (2023)89.5%자체 보고
CharXiv-R86.4%자체 보고
IPhO 202582.6%자체 보고
LiveCodeBench Pro0.80 / 3000자체 보고
SWE-Bench Verified77.4%자체 보고
Terminal-Bench 2.0Stanford × Laude Institute (2026)59.0%자체 보고
Humanity's Last Exam58.4%자체 보고
SWE-Bench ProPrinceton NLP (2024)52.4%자체 보고
ARC-AGI v242.5%자체 보고
FrontierScience Research38.3%자체 보고

Search

DeepSearchQA74.8%자체 보고

Vision

ScreenSpot Pro84.1%자체 보고
MMMU-Pro80.4%자체 보고
SimpleVQA0.71 / 100자체 보고
ERQA64.7%자체 보고
ZEROBench0.33 / 100자체 보고

AA 평가 지수

(Artificial Analysis)
Tau2(Sierra + U Toronto + Vector Institute (2025))
91.5
Gpqa(NYU + Cohere + Anthropic (2023))
88.4
Lcr(Artificial Analysis)
78.0
Ifbench(Google Research (2023))
75.9
Terminalbench V2 1
62.2
Coding Index(Artificial Analysis)
58.6
Terminalbench Hard(Stanford × Laude Institute (2026))
45.5
Hle(Center for AI Safety + Scale AI (2025))
40.7
Intelligence Index(Artificial Analysis)
31.3

LLM Stats 카테고리 점수

(LLM Stats (zeroeval))
Physics
90
Biology
90
Chemistry
90
Communication
90
Frontend Development
80
Grounding
80
Tool Calling
80
Image To Text
70
Multimodal
70
Reasoning
70
Search
70
General
70
Code
70
Vision
70
Math
60
Spatial Reasoning
60
Healthcare
60
Agents
60
Science
40

가격

입력 가격무료
출력 가격무료
혼합 가격 (3:1)무료

속도

토큰/초0.0
첫 토큰 지연0.00s
첫 응답 지연0.00s

공급자 가격 순위

프로바이더 데이터가 없습니다

외부 링크