Just now, Meta CEO Mark Zuckerberg personally announced the launch of the fourth-generation flagship model Muse Spark 1.3 in the early hours of September 3 Taiwan time. This marks the fourth version iteration in the past five months and Meta’s first entry into the top tier of AI models. According to data from independent evaluation firm Artificial Analysis, Muse Spark 1.3 ranks third globally with a score of 62 on the Intelligence Index, trailing only two Anthropic models and officially surpassing OpenAI and Google.

Five-Month Sprint: From Catching Up to Overtaking
The Muse Spark family is evolving at an astonishing pace: the 1.1 version released in July scored only 53, the 1.2 version in August improved to 57, and now Spark 1.3 has surged to 61–62. The public release, Spark 1.3 xhigh, ties with GPT-5.6 Sol (max), Grok 4.6 (high), and Claude Opus 5 (high) at 61, while the max version, available only to partners for internal testing, reaches 62—surpassed only by Claude Fable 5.1 (max) at 66 and Claude Opus 5 (max) at 63.

The implication behind this ranking is clear: Google’s top-scoring Gemini 3.8 Flash (high) only managed 59 points, placing it beyond fourth. OpenAI and xAI’s highest configurations also failed to surpass Muse Spark 1.3. Meta, which once lagged behind during the Llama 4 generation, has now secured its position among the world’s top three AI model providers.
Muse Spark 1.3 is now also available for free on the OpenCode platform, and developers on X have already reported a solid real-world experience. Some users who switched to Muse Spark 1.3 in Codex said it’s “absolutely amazing,” while other developers have paired it with Claude Code to handle complex tasks. Some users even observed on DeepSWE (Autonomous Software Engineering Benchmark) that Gemini 3.8 Flash briefly held the top spot, only to be replaced by Muse Spark 1.3 a few hours later.
Meta Muse Spark 1.3 is now free on OpenCode
— OpenCode (@opencode) September 3, 2026
The proxy mission is the biggest highlight.
According to Artificial Analysis’s detailed evaluation, this round of growth mainly comes from agentic tasks and scientific reasoning capabilities. Tau3-Bench Banking (banking agent decision-making) jumped from 35% in version 1.2 to 47% in xhigh, and the max version reached 52%, the highest score among all tested models. Terminal-Bench 2.1 (terminal operations) improved from 80% to 85%. The GDPval-AA v2 Elo score, which measures white-collar work quality, also surged from 1,615 to 1,709 (xhigh) and 1,754 (max). The reason the max version achieves higher agentic scores is that it spends more time on reasoning: it uses 62% more reasoning tokens than xhigh on GDPval-AA v2, and 28% more on Tau3-Bench Banking.

Scientific reasoning also showed progress: CritPt rose from 18% to 26%, GPQA Diamond improved from 90% to 94%, and Humanity’s Last Exam and SciCode each increased by 2 to 3 percentage points.
However, there are also areas where it has regressed. The long-text retrieval score dropped from 83% to 79%, and encyclopedia knowledge accuracy has also declined slightly. But Artificial Analysis points out that the drop in encyclopedia scores is largely because the model more often says “I’m not sure.” Its reluctance to answer randomly actually lowers the hallucination rate, which is a trade-off.
Pricing strategy: using cost-performance value to break through the market
Meta continues to maintain an aggressive pricing strategy. Muse Spark 1.3 charges $1.25 per million tokens for input (about NT$40) and $4.25 for output (about NT$138), with cache hits at just $0.15—exactly the same as version 1.2 (if using the data-sharing muse-spark-1.3-contributor, the price is even lower: just $0.1 per million tokens for input and $0.2 for output). But when calculated by cost per puzzle, xhigh only requires $0.55, making it the cheapest option among comparable models: Grok 4.6 needs $0.94, GPT-5.6 Sol needs $0.95, and Claude Opus 5 is as high as $1.23.

Efficiency gains are also a selling point. Wang said Muse Spark 1.3 can complete the same tasks with about 25% less token usage than version 1.2. The model also supports parallel processing of multiple workflows, and handles long instructions and cross-task context retention more smoothly than the previous generation. Crypto Briefing’s report noted that Muse Spark 1.3 reduced tool calls by 20%.
Additionally, the model has a clearer understanding of the boundaries of its own capabilities. Before executing irreversible operations, the model proactively requests confirmation—this safety mechanism is stricter than it was in the past.
one-million-token context and multimodal capabilities
Muse Spark 1.3 maintains an ultra-long context window of 1 million tokens, supporting multimodal input for text, images, and video. The model can be accessed via the Meta API and has also been integrated into Muse Code, Meta’s code editor positioned against Claude Code and OpenAI Codex. The OpenCode platform also announced free access to this model within hours of its release.
Zuckerberg specifically emphasized in his tweet that Muse Spark 1.3’s progress in coding and agent tasks is “the biggest leap to date,” describing its performance as “so cheap it’s almost not worth mentioning.” There are already users in the market who have tested it on the Codex platform and given positive feedback, saying its performance is comparable to top-tier models but at only one-third of the price.
Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we’ve made so far on coding and agentic work. Try it in Muse Code and our API.
Next up 🍉 and Muse Spark open weights releases coming soon. pic.twitter.com/XQQEDEJGD7
— Mark Zuckerberg (@finkd) September 2, 2026
Watermelon and Open-Source Weights: Meta’s Next Move
Zuckerberg left two cliffhangers at the end of his tweet. One is a new model codenamed 🍉 (watermelon). According to SiliconANGLE, watermelon is described as a next-generation model that is larger in scale and more powerful than the Muse Spark series, and is still in development. Wang only said that “watermelon will be extremely competitive.” Meta once enjoyed a very high reputation in the open-source model space thanks to the Llama series, but that influence is now declining as Muse Spark shifts to a commercialized, closed strategy. If it subsequently open-sources the weights of version 1.3, Meta has a chance to reclaim its leading position in the open-source camp.
The other is the upcoming “Muse Spark open-source weights.” Currently, the open-source model camp is almost entirely led by Chinese teams: Kimi K3 (max) and GLM-5.3 (max) both score 60, while Alibaba’s Qwen3.8 scores 58. If Meta releases Muse Spark 1.3-level weights, it would become the first open-source solution capable of matching closed models.
Conclusion
The release of Muse Spark 1.3 marks Meta’s official pivot from the open-source strategy of the Llama era to commercial closed models, proving within five months that this path is viable. With five iterations, each showing clear improvement, yet pricing squeezed to one-third of competitors’, this pace has taken Meta from a chaser to the “most price-competitive” player in the leading pack. The upcoming final decisions on the watermelon model and open-source weights will determine Meta’s long-term positioning in the AI race. For developers, Muse Spark 1.3 is already available on the Meta API and Muse Code, ready for immediate trial.