Meta Muse Spark Model

Muse Spark

Muse Spark is Meta Superintelligence Labs' first Muse-family model, built for native multimodal reasoning, tool use, visual chain of thought, and multi-agent orchestration. Use this page to compare its official demos, benchmark signals, scaling data, and MyClaw workflow fit.

Apr 8, 2026
Official Announcement
58%
Humanity's Last Exam with Tools
38%
FrontierScience Research
Capabilities

Where Muse Spark Fits in Agent Workflows

Use Muse Spark when an agent needs to reason over images, tools, health context, and parallel reasoning paths in one hosted workspace.

Visual Reasoning Over Real Context

Muse Spark is designed to integrate visual information across tools and domains, including visual STEM questions, entity recognition, and localization.

Visual Chain of Thought

Meta highlights visual reasoning traces and dynamic annotations, which help the model explain or troubleshoot what it sees instead of only returning text.

Tool Use and Agent Orchestration

Muse Spark supports tool use and multi-agent orchestration, making it more relevant for agent workflows than a plain chat-only model.

Health and Wellness Explanations

Meta says it worked with over 1,000 physicians to curate training data for more factual and comprehensive health responses.

Contemplating Mode for Hard Tasks

For difficult reasoning, Contemplating mode coordinates multiple agents in parallel rather than relying on one longer single-agent pass.

Limited API Access Today

Muse Spark is available in Meta AI and the Meta AI app, while private API preview access is limited to select users. MyClaw tracks the model as availability expands.

Official Demos

Official Muse Spark Demos from Meta

These demos show the clearest user-facing shape of Muse Spark: visual annotation, health explanation, and movement analysis powered by multimodal reasoning.

Dynamic Visual Annotations

Muse Spark can annotate objects and steps in a scene, making it useful for troubleshooting, instruction following, and visual explanations.

Health and Nutrition Displays

Meta shows Muse Spark turning visual context into interactive health and nutrition displays that are easier to inspect than plain text.

Fitness Form Breakdown

Muse Spark can compare movement, mark body positions, and explain how a pose or exercise form can improve.

Muse Spark Benchmark Signals from Meta

After the demos, use Meta's benchmark chart as the evidence layer: it reports performance across reasoning, science, multimodal perception, health, and agentic tasks.

Official Muse Spark Contemplating Mode Benchmark Chart
Scaling Evidence

The Research Signals Behind Muse Spark Scaling

Meta frames Muse Spark progress around three levers: more efficient pretraining, smoother reinforcement learning gains, and more efficient test-time reasoning.

01
Pretraining

More Capability per Unit of Compute

Meta says architecture, optimization, and data curation changes helped Muse Spark extract more capability from training compute.

The official scaling-law chart compares Muse Spark with Llama 4 Maverick and leading base models, supporting Meta's claim that the new recipe is significantly more compute efficient.

Official Muse Spark Pretraining Scaling Chart
02
Reinforcement Learning

Predictable Gains from RL Compute

Meta says its RL stack produces smoother improvements, even though large-scale reinforcement learning can be unstable.

The training plot shows pass@1 and pass@16 rising with RL steps, while the held-out evaluation chart suggests those gains generalize beyond tasks seen during training.

Official Muse Spark Reinforcement Learning Scaling Chart
03
Test-Time Reasoning

More Reasoning Without the Same Latency Hit

Muse Spark combines thought compression with parallel agent orchestration to spend reasoning budget more efficiently.

Meta says thinking-time penalties can reduce unnecessary reasoning tokens, while multi-agent thinking can improve performance at comparable latency.

Official Muse Spark Test-Time Reasoning Scaling Chart
Safety

Safety and Deployment Checks

Meta evaluated Muse Spark across frontier risk categories, alignment behavior, adversarial robustness, and evaluation awareness before deployment.

Frontier Risk Categories

Meta reports strong refusal behavior across high-risk biological and chemical domains and says Muse Spark stayed within safe margins for the measured frontier risk categories.

Official Muse Spark Safety Evaluation Chart

Evaluation Awareness

Apollo Research found high evaluation awareness on a near-launch checkpoint. Meta says follow-up work did not make this a blocking release concern, but it remains an area for research.

Official Muse Spark Evaluation Awareness Chart

Muse Spark FAQs

Meta Muse Spark

Run Muse Spark in a Hosted MyClaw Agent

Use MyClaw to evaluate Muse Spark-style multimodal workflows with files, tools, browser-ready execution, and a managed agent workspace as model access expands.