• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - Latest Technology News - M5 Ultra AI Performance Tested: Apple Wasn’t Kidding—Local AI Performance Gains Match the Marketing Claims

M5 Ultra AI Performance Tested: Apple Wasn’t Kidding—Local AI Performance Gains Match the Marketing Claims

KOCPC Editor by KOCPC Editor
September 24, 2026
in Latest Technology News

Apple announced in August a new Mac Studio equipped with the M5 Ultra, claiming local AI performance up to 4.3 times faster, storage speed 2 times faster, CPU speed 1.3 times faster, and unified memory bandwidth reaching 1.2TB per second. Well-known overseas tech YouTuber Alex Ziskind put the M5 Ultra head-to-head against the previous-generation M3 Ultra, and the conclusion in his title was blunt: “Apple Wasn’t Messing Around.” This time he tested the fully maxed-out M5 Ultra version: 80-core GPU, 256GB unified memory (the 512GB version goes on sale the following month), 8TB SSD, with the whole machine priced at $14,299 (about NT$465,000); the M5 Ultra version of the Mac Studio starts at $5,499 (about NT$179,000). He joked that this generation had “a performance upgrade, and quite a price upgrade too.” The test machine was loaned by Apple, while the comparison M3 Ultra was bought with his own money.

M5 Ultra AI performance tested: local AI performance gains match marketing claims

Test configuration: 36 cores, no efficiency cores

The M5 Ultra CPU has 36 cores in total, made up of 24 performance cores and 12 super cores, with no efficiency cores at all in the Mac Studio. To keep things fair, both machines ran the same builds of llama.cpp and MLX and loaded exactly the same model files. For video processing, the M5 Ultra’s media engine has twice as many encode/decode blocks as the M5 Max, and it is officially claimed to play up to 33 simultaneous 8K ProRes 422 30fps video streams. Geekbench 7 benchmark results also show that the 36-core M5 Ultra scores 3,774 single-core and 52,516 multi-core, about 30% and 35% higher than the average for the 32-core M3 Ultra, roughly in line with the official claim of “1.25x single-threaded and 1.3x multi-threaded.”

Developer Workloads: Compilation and Storage Take Off

For developers, this machine is “overkill.” The Speedometer web performance test went from 47 points on the M3 Ultra to 60.2 points. Real-world compile tests are even more telling: for a large .NET project containing 100,000 namespaces and classes, compile time dropped from 71.7 seconds to 52.4 seconds; a pure Python Mandelbrot algorithm test dropped from 10.25 seconds to 5.8 seconds, a full two times faster, prompting the host to exclaim, “I’ve never seen anything this fast.”


The official claim is that storage is 2x faster, but real-world tests go even beyond that. AmorphousDiskMark sequential read/write: M3 Ultra reads at 5930MB/s and writes at 7788MB/s; M5 Ultra reads at 14,888MB/s and writes at nearly 20,000MB/s. Random read/write, which compilation relies on even more, is also more than twice as fast; his comment is, “The marketing department’s 2x labeling is too conservative.” This also has a direct impact on AI: when loading a 140GB model file, read speeds are much faster.

Two key numbers for local LLMs

Testing local large models comes down to just two numbers: prompt processing (PP) and token generation (TG). PP is the process of the model “reading” the prompt, and it consumes GPU compute; TG is the process of the model “writing” the answer, and it consumes memory bandwidth.

The architectural upgrade of the M5 Ultra is exactly what’s needed: the GPU cores in the M3 generation were just ordinary display cores, while each GPU core in the M5 generation has a built-in Neural Accelerator, meaning every core has dedicated matrix multiplication AI hardware. The “has tensor” flag when llama.cpp starts shows false on M3 Ultra and true on M5 Ultra, meaning the software has enabled the Neural Accelerator path for the M5 family. Apple also officially emphasizes that this is the first time a Neural Accelerator has come to an Ultra-class chip, delivering up to 4.3x peak AI compute compared with M3 Ultra.

Memory bandwidth measured: 1.2TB/s achieved, about 86%

The official rating is 1.2TB/s. First, run the venerable STREAM benchmark (a CPU-side test): M3 Ultra at 312.9GB/s versus M5 Ultra at 583.6GB/s. Switching to a GPU-side microbenchmark (which is what actually matters for AI): M3 Ultra measured 724GB/s (officially rated 819GB/s, hitting about 90%), M5 Ultra measured 1,039GB/s (officially rated 1.2TB/s, hitting about 85% to 86%), with measured bandwidth 1.4 times that of M3 Ultra.

Three models tested in practice: both MoE and dense models validated.

  • DeepSeek V4 Flash (284 billion parameters, MoE): Token generation increased from 37 to 53 tokens/s (about 1.5x), consistent with the bandwidth increase; prompt processing increased from 483 to 1,485 tokens/s (3x).
  • GPT-OSS 120B (120 billion parameters, MoE)Token generation: 90 vs. 130 tokens/s; prompt processing: 1,300 vs. about 3,200 tokens/s. In long-context tests, after already having a 64,000-token context, feeding another 32,000 tokens took the M3 Ultra 80 seconds and the M5 Ultra only 56 seconds.
  • Qwen 3 8.27B (dense model): Token generation 37 vs. 54 tokens/s (1.47x); prompt processing 430 vs. 1,800 tokens/s, over 4x. Separately, tested with a real 14,000-token code file, prompt processing dropped from 33 seconds to 8 seconds, a measured 4.2x, delivering on the official promise of “up to 4x.”

MoE (mixture of experts) and dense models both land at 1.4 to 1.5 times in token generation, exactly matching the measured 1.4 times memory bandwidth, validating the theory that TG (token generation) is memory-bandwidth-bound. The biggest generational difference is in prompt processing, which has long been the weak spot of the M3 and M4 generations; the neural accelerator makes up for it in one go.

The premise of 4x and the power consumption cost

Don’t rush to swipe your card when you see this—4x prompt processing has a catch: the prompt has to be long enough. You need about 14,000 tokens to reach the 4x threshold; a short 350-token pure-chat prompt only gets 1.6x. “Users who mainly chat won’t notice much, but feed it a few source code files and they will.”

Efficiency gains are limited: during generation, chip power consumption is about 45 W, versus 36 W for the M3 Ultra (both GPU plus CPU, not wall power), and per-watt token efficiency is only about 12% higher. Idle power is lower, however: the M5 Ultra is about 8 to 10 W, while the M3 Ultra is 11 to 12 W. But under full GPU load, the difference is stark: the M3 Ultra approaches 200 W, while the M5 Ultra exceeds 400 W, the fan noise is noticeably louder, and the chassis temperature is around 43–44 degrees, “considerably hotter to the touch than the M3 generation.”


Summary: 1.5x token generation, 2 to 4x prompt processing

The host’s conclusion: 1.5x token generation is a “nice upgrade,” and 2x to 4x prompt processing will “definitely be noticeable” to coding agent users. In the previewed multitasking test, when processing 8 requests simultaneously, M5 Ultra output 134 tokens/s, while M3 Ultra was 75 tokens/s. He also previewed that follow-up coverage will compare MLX and llama.cpp, 8-bit models in depth, and test larger models with a 512GB machine and a Mac Studio cluster.

For more details, you can watch the video and turn on CC subtitles yourself:

 

Source: KOCPC Chinese

Tags: M3 UltraM5 UltraMac Studio

Recent Posts

  • Apple secretly developing a screenless fitness tracker to take on Whoop: fabric band plus sensor module could arrive as early as 2028.
  • From NT$2,000 to NT$18 million in annual revenue for a physical store! TikTok lights up heartwarming Mid-Autumn business opportunities as baking brands and local small farmers use short videos to unlock a new business formula.
  • M5 Ultra AI Performance Tested: Apple Wasn’t Kidding—Local AI Performance Gains Match the Marketing Claims
  • 10,000mAh-Class Battery Life! The OPPO A7 Pro Series Introduces a High-Spec Pro Max Version for the First Time, Launching Across Taiwan on October 1.
  • vivo V80 Lite lands in Taiwan, goes on sale 10/1: 10,000mAh BlueSea battery sets a Guinness World Record, starting at NT$16,990

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology