• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - You can run a 120B large model without 60GB of memory! An overseas user successfully linked a phone, Mac, and Windows to achieve it.

You can run a 120B large model without 60GB of memory! An overseas user successfully linked a phone, Mac, and Windows to achieve it.

Rocky by Rocky
September 29, 2026
in AI Trends and Related News

Generally speaking, to run a 120-billion-parameter large language model locally, the standard approach is to first get an 80GB data center GPU, or buy a Mac Studio with 128GB unified memory or an AMD Ryzen AI Max+ mini PC, which costs at least over NT$100,000, and most people can’t do it. Recently, however, an internet user abroad successfully ran OpenAI’s open-weight gpt-oss-120b model without spending money on new hardware by chaining together phones, laptops, and a Mac mini at home.

Of course, just because it can run doesn’t mean it runs smoothly; the generation speed is a dismal 1.1 tokens per second.

A phone plus five computers form an AI cluster! It can run gpt-oss-120b, but only at 1.1 tokens per second.

Recently on the r/LocalLLM subreddit, a user named Medicine_Blogscanner shared a way to run a 4-bit quantized version of gpt-oss-120b split across six consumer-grade devices, with memory usage down to 47GB and a generation speed of 1.1 tokens per second.

According to this netizen’s original post, the six devices are: a Windows laptop with 12GB of RAM, a mini PC with an RTX 3060 (12GB of VRAM), a 16GB Mac mini, a 16GB M3 MacBook, a 2017 Intel MacBook that can only run on CPU, and his Android phone. The phone is a Galaxy S24+. Half of the six devices are connected via wired Ethernet, and the other half are connected via Wi-Fi.

He also wrote:

老實說,我真的沒想到這會成功。這幾台機器沒有一台能自己載入 120B 的模型,差得遠了。所以我乾脆把它們全部丟進一個叢集裡,硬試看看。

Ran a 120B model across 6 computing devices that had no business running it!
byu/Medicine_Blogscanner inLocalLLM

And the whole setup process goes roughly like this. He set up the mini PC with an RTX 3060 as the host machine, then let the RAMDeck software handle the allocation on its own. The host machine first keeps the portion it can accommodate locally, then distributes the model piece by piece according to the actual capabilities of the other devices.

The allocation order is: the GPU fills up first, then two Macs accelerated with Apple’s own GPU interface, Metal, then devices that can only run on CPU, and finally the phone also gets a small slice. Interestingly, of the six devices, only five actually handle the model; the 2017 Intel MacBook gets nothing at all, and he agrees with this allocation because this MacBook is simply too old.

From the screenshot, a single Tiny handles about 60%, and of the 28.7GB allocated to it, only 9.7GB is in the RTX 3060’s video memory, while the other 19GB is in system memory, with the CPU handling computation. The Windows laptop was allocated 7.9GB, more than the Mac mini’s 7.4GB. The M3 MacBook was allocated 2GB, and the Galaxy S24+ only got 1.19GB:

If the screenshot is unclear, you can refer to the table below:

Node name (inferred corresponding device) GPU allocation CPU allocation Total Proportion of the total
Tiny (RTX 3060 mini PC, host machine) 9,675MB 19,040MB 28,715MB Approximately 61%
ACER-GAU (12GB Windows laptop) — 7,940MB 7,940MB Approximately 17%
Gaurangs-Mini(16GB Mac mini) 7,473MB — 7,473MB Approximately 16%
Anujas-MBP(16GB M3 MacBook) — 2,013MB 2,013MB About 4%
SM-S926U1(Galaxy S24+) — 1,190MB 1,190MB Approximately 2.5%
2017 Intel MacBook — — Unallocated 0%
Total 17,148MB 30,183MB 47,331MB 100%

Although gpt-oss-120b ran successfully, the loading process was not smooth. A full load took about 11 minutes, with most of the time spent waiting for the model’s chunks to slowly transfer over Wi-Fi to a slower device. It took him three tries to succeed, because the program had a hard-coded 10-minute timeout limit, so it was interrupted before loading finished.

In the post, he also gave others a piece of advice: the load timeout should be adjusted according to the model size, not hardcoded.

The final measured speed was 1.1 tok/s, and some people may have no idea what that number means.


A token is the basic unit of text processed by a language model. In English, on average, 1 token is roughly equivalent to 0.75 English words. In other words, 20 English words usually amount to about 25 to 30 tokens. At a generation speed of 1.1 tokens per second, a sentence of about 20 English words might take over 20 seconds to fully appear. If you ask it to generate a long response of 500 tokens, it will take about 7 minutes and 35 seconds—and that’s not even counting the time the model spends processing the input prompt.

It’s normal for it to be this slow. For every token generated, the data has to pass through each device in sequence over the local network, and only after it finishes computing can the next one start, so the speed is bottlenecked by the slowest device.

He also recorded the whole process as a video and posted it on YouTube:

Source: KOCPC Chinese

Tags: gpt-oss-120bLLMLocal AIOPENAIRAMDeck

Recent Posts

  • 5 Common Bluetooth Myths: These beliefs you’ve always held may not actually be true
  • You can run a 120B large model without 60GB of memory! An overseas user successfully linked a phone, Mac, and Windows to achieve it.
  • How do Apple Watches made of aluminum alloy, ceramic, and titanium differ?
  • Qwen4 Internal Test Results Leaked Early? Frontend Rumored to Reach Opus 5.5 Level; Qwen4 27B Rumored to Be Open-Sourced Early in October
  • MUSE got too aggressive with selling and screwed up! An AI managing a netizen’s online storefront gave their address to a buyer for an in-person meetup within a day, and even replied, “I’m at home.”

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology