What Meta announced

Meta released Muse Spark on April 8, a year after Llama 4 and the first model out of Meta Superintelligence Labs, the reorganised division that Alexandr Wang runs. It is live in the meta.ai chat interface and the Meta AI app, with a private API preview for selected users. There are two modes, Instant and Thinking, and a third called Contemplating is promised. There are no weights. Meta says it plans to open-source future versions, and Wang's own framing is that this is step one and that there are rough edges in model behaviour the team will polish over time.

The claim that got the most attention is the efficiency one. Meta says Muse Spark reaches the capabilities of Llama 4 Maverick with over an order of magnitude less compute, and that it is significantly more efficient than the leading base models available for comparison. On self-reported benchmarks Meta places it alongside Claude Opus 4.6, Gemini 3.1 Pro and GPT-5.4 on selected tests, and behind them on Terminal-Bench 2.0.

What an order of magnitude less compute can mean

The efficiency claim is the one worth reading slowly, because it is compatible with several very different facts about the model. The clean interpretation is that Muse Spark was pretrained with a tenth of Maverick's training FLOPs and matches it on evaluations, which would be a real algorithmic or data result. But the same sentence is true if the comparison is inference compute, which would be a statement about a smaller active parameter count or a better serving stack. It is also true if it describes total compute to reach Maverick's score on a chosen set of evals, where the choice of evals does a lot of the work.

Meta has not said which of these it means, and the comparison point matters. Llama 4 Maverick is a year old, so matching it cheaply is a lower bar than matching the frontier cheaply. The interesting version of the claim would be a compute comparison against Opus 4.6 or GPT-5.4, and that comparison is precisely the one that is not made. Until Meta publishes a training compute figure, the honest description is that Muse Spark is a much more efficient model than Meta's previous one, and that its distance from the frontier is measured on benchmarks rather than on FLOPs.

The independent numbers

Artificial Analysis scored the model at 52 on its intelligence index, which puts it fifth, behind Gemini 3.1 Pro Preview, GPT-5.4 and Claude Opus 4.6. The Decoder's write-up pulls out two contrasting figures from that evaluation. In Thinking mode Muse Spark scores 50.2 on Humanity's Last Exam without tools, ahead of Gemini 3.1 and GPT-5.4 Pro on that test. On the GDPval-AA work tasks it scores 1,427 points, behind Claude Sonnet 4.6 at 1,648 and GPT-5.4 at 1,676.

That split is a familiar shape. A model that does well on a hard academic exam and less well on tasks that look like real work usually has strong reasoning post-training and weaker agentic and tool-use training, which fits both the Terminal-Bench gap and Wang's line about rough edges. Simon Willison counted 16 tools wired into the meta.ai interface, including web search, Python execution, image generation and visual grounding, so the product layer is rich. His pelican-on-a-bicycle test found Thinking mode much better than Instant, which is what you would expect and is not evidence of much beyond the modes doing what they say.

The company that made open weights normal

The part of this launch that will be remembered has nothing to do with the benchmarks. Meta is the company whose Llama releases made open weights a mainstream option for everyone from startups to national labs, and it had never before led a generation with a model it did not release. Muse Spark cannot be downloaded or run locally, and no explanation for that choice has been offered beyond the promise that later versions will be opened.

For research groups like ours the effect is concrete. Llama 4 was the last Meta model we could put under an interpretability tool, and a private API preview is not a substitute. If the efficiency claim is real, it is also exactly the kind of result that the community could have verified and built on with the weights in hand, and that cannot be verified from a chat window. A closed model with an open-weights promise attached is a closed model.

We are reluctant to read the switch as a settled strategy. Wang's statement that future versions will be open-sourced could be a sequencing choice, releasing the closed version first to protect a product launch and opening a smaller or older variant later. That was not how Llama worked, and the difference is the point.

What would settle the efficiency question

Three numbers would turn the announcement into a result. The total pretraining compute for Muse Spark and for Maverick, in FLOPs. The active and total parameter counts. And the same benchmark suite run by an outside party at fixed decoding settings, with a compute-matched comparison against at least one open model of known size. None of these require weights. A technical report would do.

If Meta publishes them, the order-of-magnitude claim becomes something the field can learn from. If it does not, then Muse Spark joins the list of closed models whose efficiency we have to take on faith, and the lab that once let us check its work has stopped doing so. We would rather be wrong about that.

Sources

  1. Simon Willison, Muse Spark (April 8, 2026)
  2. The Decoder, Meta's Muse Spark is its first frontier model and its first without open weights