As everyone knows, the M5 Ultra and DGX Spark are currently the most hyped consumer local AI computers, so many people are probably curious about which of the two is stronger for local applications. However, the M5 Ultra’s current top configuration is 256GB RAM (512GB goes on sale at the end of October), while a single DGX Spark has only 128GB RAM, so a direct comparison isn’t really fair. Recently, overseas AI YouTuber Alex Ziskind put a Mac Studio (M5 Ultra, 256GB) up against a dual-unit DGX Spark cluster (128GB each x 2), and the results might be more surprising than you’d expect.

M5 Ultra vs DGX Spark: What the Two 256GB Models Look Like
The two have the same amount of memory on paper, but they look completely different. This Mac Studio is a maxed-out M5 Ultra with 256GB of unified memory and an 8TB SSD; the configuration he has costs $14,000, or about NT$455,000.

On the other side are two NVIDIA DGX Sparks, each with 128GB plus 4TB of storage. At launch they were $3,999 each; the official price has now risen to $4,699 (about NT$153,000). Together, the two cost about NT$305,000, plus a not-so-cheap 200Gbps QSFP cable.

That link runs RoCE (RDMA over Converged Ethernet), allowing the two machines to directly read and write each other’s memory; the model is split using vLLM’s tensor parallelism (Tensor Parallel), with each layer split in half across the two machines. Every token generated requires the intermediate results to be exchanged over the link before the next layer can begin (in fact, Apple has similar technology too). RDMA-over-Thunderboltusing TB5 cables). Why must there be two machines? DeepSeek V4 Flash is about 150 to 160GB after quantization, so it cannot fit into a single 128GB Spark; when split across two machines, each uses about 111GB, leaving just enough room to operate. A Mac, by contrast, packs the whole thing into a single memory pool. The models he tested were DeepSeek V4 Flash and Qwen 3.8 Flash Next, with both running 4-bit quantized versions.

Generation speed is tied; memory bandwidth decides the winner.
First, look at “writing,” i.e., token generation. For DeepSeek, both sides are around 38 tokens per second, a tie; for Qwen, the Mac’s 45 per second versus Spark’s 38 gives the Mac a win. The reason points to memory bandwidth: the M5 Ultra’s official spec is 1.2 TB/s, while each Spark has only 273 GB/s; even two combined are still less than half of the Mac’s, and he measured that nominally 200 Gbps link at about 111 Gbps, with synchronization overhead eating up some more.

Software also has a big impact. The same DeepSeek on a Mac generates about 40 tokens per second with Llama.cpp; switching to MLX jumps straight to 53, a 34% speedup from just changing the engine; when both use Llama.cpp, the difference is within 2%.

Reading prompts is Spark’s home turf.
Now look at “read,” i.e. prompt processing: the picture is completely reversed. At 2,000 tokens, both sides are so fast you don’t even notice; at 8,000 tokens, DeepSeek takes about 4 seconds on Spark and about 10 seconds on Mac; at 16,000 tokens, the Mac wait doubles to 21 seconds.

The gap widened in real-codebase TTFT (time to first token) tests: with 32,000 tokens across 14 Python files fed in, dual Spark finished reading and began emitting tokens in 17 seconds, while the Mac was still reading the prompt at 50 seconds, roughly a 3x gap. The same ratio holds for a shorter Qwen prompt: Mac 26 seconds vs. Spark’s 11.2 seconds, 2.4x. Spark scaled all the way to 128,000 tokens in only 72 seconds; at that level, he simply didn’t test the Mac, and extrapolating from its 32K speed suggests a wait of over 3 minutes. This matches hardware-architecture expectations: prompt processing is bound by GPU parallel compute, and two Sparks’ combined compute beats a single M5 Ultra; token generation is bound by memory bandwidth, so the single-machine unified-memory Mac actually has the edge.


Cache saved the Mac
But real-world conversations don’t start from zero every time. The server keeps what it has already read (prefix caching). In a 16,000-token context: the first question takes 21 seconds on Mac and 8.3 seconds on Spark; follow-up questions take only 3 seconds on Mac and 1.4 seconds on Spark. That “reading a long document” cost is paid only once per session. This is how code tools work: the Agent sends the codebase once, the server keeps it, and follow-up questions only add the new parts.

The pain point is that the context keeps changing: switch repos, drop in large files, or let a conversation exceed the cache limit, and it has to be reread. Another detail he observed is that switching models also triggers a full recomputation, effectively paying the tuition all over again. This also explains the interest in the field in “disaggregated prefill/decode”: letting a compute-heavy Spark focus on reading prompts and a high-bandwidth Mac focus on generation, playing to each one’s strengths. He says people are already working on this kind of project, but it is still early-stage experimentation.

Multi-user use is the real watershed.
When multiple people use them at the same time, the difference in character between the two machines is stretched to the max. When DeepSeek is used by 4 people simultaneously on a Mac, total throughput is 66 tokens per second, dropping to 46 with 8 people; dual Spark climbs all the way to 70. Given 8,000 tokens of conversation history per person and 8 people online at once: Mac total throughput falls to only 11 tokens per second, while Spark has 25. What hurts more is the wait: on Mac, each person has to wait more than a minute to see the first token, then gets about 5 tokens per second; with Spark, they wait about 24 seconds, then get about 7 tokens per second. Simply put: a Mac is a one-to-two-person machine, dual Spark suits a small team of around four, and eight people is pushing it.

Heat, electricity, and hassle level.
At full load, the M5 Ultra draws 434 watts and the Spark cluster 410 watts—almost the same, and both would trip the power meter. On the chassis surface, the hottest point on the Mac is 49 to 50°C, while the Spark is about 48°C; the rear exhaust vent on the Mac Studio is 56°C, and inside the Spark he measured 58 to 63°C. Both are noisy and hot. He specifically added that these full-load figures do not represent standby, and when idle


The gap in setup cost is what truly separates them in daily use: Mac has almost zero setup and can even moonlight for video editing and rendering; a Spark cluster requires aligning the OS, kernel, drivers, and firmware of two machines, then installing vLLM multi-node setup and a bunch of NCCL and RoCE environment variables, and finally checking RDMA traffic. Simply put, aside from AI, Mac Studio can easily and enjoyably handle all daily work and entertainment, while Spark is specifically for AI, and you also have to wrestle with the Linux environment yourself—so it depends on your choice. Those interested can watch the introduction in the original video; it has CC subtitles:
Conclusion: Two 256GB boxes do not equal one 256GB box.
His answer is: it can fit the same model, but reads much faster. As for which to buy, first measure your workload: How large are your prompts? How long do the answers need to be? Large inputs (codebases, long documents) depend on Spark’s read speed, while long outputs and heavy single-user use depend on the Mac’s bandwidth and convenience. Don’t forget the bill: two Sparks plus cables comes to under NT$300,000, while his 8TB M5 Ultra costs NT$455,000 (actually, if you don’t buy an SSD that large, it’s around NT$350,000; the price difference is mainly the SSD that costs more than gold), and the price gap is enough to buy a high-end laptop. If your work is throwing the entire codebase at the model all day long, the Spark cluster’s read speed will directly change the experience; if it’s mostly you having long conversations with the model, the Mac’s convenience and quietness are more practical.
Source: KOCPC Chinese