More Capability per Unit of Compute
Meta says architecture, optimization, and data curation changes helped Muse Spark extract more capability from training compute.
The official scaling-law chart compares Muse Spark with Llama 4 Maverick and leading base models, supporting Meta's claim that the new recipe is significantly more compute efficient.





