• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - Latest Technology News - First oMLX performance results for M5 Ultra Mac Studio leak: prefill speed 3-4x faster than M3 Ultra, data taken down after exposure.

First oMLX performance results for M5 Ultra Mac Studio leak: prefill speed 3-4x faster than M3 Ultra, data taken down after exposure.

KOCPC Editor by KOCPC Editor
September 19, 2026
in Latest Technology News

Apple’s next-generation Mac Studio (equipped with the M5 Ultra) has not officially shipped yet, but the first batch of M5 Ultra local AI performance data has quietly appeared on the oMLX performance database, and compared with the previous generation’s specs, it looks quite impressive. The data was removed shortly after it surfaced, adding another layer of intrigue to this “early leak.”
Apple 突襲發表 Mac mini M6 與 M5 Pro:首搭 2 奈米晶片,AI 效能提升 4 倍,台灣 29,900 元起 - 電腦王阿達

According toOfficial Apple specificationsThe M5 Ultra features the first four-die UltraFusion architecture, offering up to a 36-core CPU (12 super cores + 24 performance cores), an 80-core GPU, and, for the first time, a neural network accelerator built into an Ultra chip, with peak AI compute performance up to 4.3x that of the M3 Ultra and 9.8x that of the M1 Ultra. Paired with up to 512GB of unified memory and 1.2TB/s of memory bandwidth (a 50% increase over the previous generation), this machine was positioned from its debut as a desktop computer that can run large language models entirely locally.
Apple 突襲發表 Mac mini M6 與 M5 Pro:首搭 2 奈米晶片,AI 效能提升 4 倍,台灣 29,900 元起 - 電腦王阿達

The leaked data has attracted attention for two reasons. First, this is the first actual performance result of the M5 Ultra chip running in the MLX ecosystem, rather than theoretical numbers from official slides; second, the data appeared through oMLX’s community benchmark page, meaning the testing environment, context length, and batch size all follow standardized conditions, making it easier to compare directly with other chips than ordinary individual benchmark runs.

The first batch of benchmark data leaked from oMLX

On the oMLX Community Benchmarks page (now removed), benchmark data labeled “M5 Ultra (80c) / 256GB” appeared. oMLX is an inference engine built on top of the Apple MLX framework, providing a unified community benchmarking platform that lets different Apple Silicon chips be compared under the same test conditions (same context length, same batch size). According to the leaked screenshot, the platform’s metrics are divided into two columns, PP (prefill, prompt processing speed) and TG (text generation, text generation speed), and it supports filtering by chip, model, quantization, context length, and other conditions.


According to leaked data, the performance of M5 Ultra under different model and quantization combinations is as follows:

  • Qwen3.8-27B (Q4 quantized, 8K context)prefill (prompt processing) 1,800 tokens/s, decode (text generation) 51 tokens/s
  • Qwen3.8-Flash-Next (Q4 quantization + MTP, 64K context):prefill 2,574 tokens/s、decode 80 tokens/s
  • GLM 5.3 Flash (oQ4e quantization, 16K context):prefill 955.4 tokens/s、decode 33.0 tokens/s
  • Qwen3.8-Flash-Next (Q4 quantization + MTP, 32K context):prefill 2,661 tokens/s、decode 23.8 tokens/s
  • Qwen3.8-Flash-Next (Q4 quantization + MTP, 16K context):prefill 2,681 tokens/s、decode 61.3 tokens/s

Among them, Qwen3.8-Flash-Next achieved a peak prefill speed of 2,681 tokens/s at a 16K context, and maintained 2,574 tokens/s at a 64K context, showing that the M5 Ultra’s prompt processing capability barely degrades noticeably with context length.

Another developer also saved the relevant values, which can serve as a supplement:

The key to generational upgrades: prefill is 3-4x faster

Compared with the M3 Ultra data also obtained by the community, the most noticeable aspect of this upgrade is prefill speed. In the MLX official community’s systematic M3 Ultra tests, Qwen 32B Q4 quantization had a decode speed of only about 27.5 tokens/s at 8K context, while Q2 quantization only reached about 47.6 tokens/s at 1K context. The tests at the time also noted that M3 Ultra’s prefill was limited by a compute bottleneck (about 54 TFLOPS FP16), regardless of quantization level.

The AI community has observed that, compared with equivalent M3 Ultra configurations, the M5 Ultra delivers about 3 to 4 times faster prefill speed, and it is considered one of the biggest generational improvements of this chip generation. This is in line with Apple’s official claim that “AI peak compute performance is 4.3 times that of the M3 Ultra.”

In terms of decode speed, M5 Ultra’s performance of about 50-80 tokens/s on a 27B model also shows a clear increase compared with M3 Ultra’s data on 30B-class models, but because the two use different models and quantization methods, the numbers can only serve as a directional reference and should not be directly equated. Based on experience from the MLX community, the bottleneck in the decode stage is mainly memory bandwidth rather than compute; the benefit of M5 Ultra’s 1.2TB/s bandwidth to text generation speed, more than prefill, will be directly reflected in long-context and long-output application scenarios.

The data was removed shortly after being exposed.

After this batch of data sparked discussion online, the M5 Ultra project on oMLX has disappeared. In a follow-up tweet, the poster said they could not confirm whether the original tester deleted it themselves or whether the oMLX platform took it down afterward.

Some of the first M5 Ultra Studio performance numbers have been posted on oMLX and they are looking very good.

Qwen3.8-27B at Q4:
1,800 tokens per second prefill
51 tokens per second decode at 8k context

Qwen3.8-Flash-Next at Q4 MTP:
2,574 tokens per second prefill
80 tokens… pic.twitter.com/UgN3z6VQXf

— Mike Bradley (@MikeBradleyAI) September 17, 2026

The key issue is that the M5 Ultra Mac Studio will not officially begin shipping to general consumers until September 22, meaning whoever submitted this batch of data got the machine at least 5 days before the official release—possibly media, developers, or insiders who received a sample unit early. This also explains why the data was removed: before a product officially goes on sale, unofficial performance data usually involves confidentiality agreements or product launch timing considerations.

In the Taiwan market, the Mac Studio with M5 Ultra has been available for pre-order since August 27, starting at NT$199,900 (education pricing starting at NT$185,390), and the version with 512GB unified memory is expected to launch in late October.

Apple 突襲發表 Mac mini M6 與 M5 Pro:首搭 2 奈米晶片,AI 效能提升 4 倍,台灣 29,900 元起

The Significance of Local AI Deployment

For the developer community, this data symbolizes that the threshold for “flagship-grade local AI” has been lowered another notch. Previously, to smoothly run 70B-class or larger models on a Mac, one often had to endure slow text generation or noticeable slowdowns with long contexts. The M5 Ultra is positioned differently: an 80-core GPU plus up to 512GB of unified memory can simultaneously accommodate multiple 30B-class models, and a prefill speed of 2,500+ tokens/s makes the scenario of “throwing a large amount of documents in for analysis” practical. Apple has also officially mentioned that multiple Mac Studios can be linked into a cluster via Thunderbolt 5 and RDMA, and a cluster of four can achieve AI inference speeds up to 3 times that of a single unit, further expanding the upper limit of local inference scale.

For developers and AI researchers in the Taiwan market, the price range of the M5 Ultra Mac Studio (starting at NT$199,900) is still in the professional workstation class, but given that it can run large models locally without relying on the cloud at all, it does have a certain appeal compared with the long-term cost of renting cloud GPUs. However, the 512GB memory version will not launch until the end of October, so users who really want to tackle the largest-scale models will have to wait another month.

Summary

Although the leaked data only covered the 80-core GPU + 256GB memory configuration and has already been taken down, the direction is quite clear: the performance leap of the M5 Ultra in the prefill stage is the most anticipated part of this generation of chips, and the officially claimed “4.3x improvement in AI performance” is not just a number on paper. For developers planning to run large language models on a Mac, or even to link multiple Mac Studios via Thunderbolt 5 to form a cluster, real benchmark data will soon begin to emerge after shipments start on September 22.

Source: KOCPC Chinese

Recent Posts

  • Meta Muse Lands on Mac: Personal AI Agent Begins Operating Your Files, Messages, and Calendar (Muse for Mac)
  • First oMLX performance results for M5 Ultra Mac Studio leak: prefill speed 3-4x faster than M3 Ultra, data taken down after exposure.
  • After the iPhone 17, the frame material was changed back to aluminum alloy. Is it better than titanium?
  • The iPhone 18 Pro series, Apple Watch Series 12, Apple Watch Ultra 4, and AirPods 5 are now on sale—a first look, with warm “Burgundy Red” and striking “Glacier Blue.”
  • The battery replacement cost for the iPhone 18 Pro / Pro Max has gone up again! Compared with 5 years ago, the increase has doubled.

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology