• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - An overseas YouTuber tested running a local LLM on a 20-year-old Pentium 4 PC; answering a single question took 33 minutes.

An overseas YouTuber tested running a local LLM on a 20-year-old Pentium 4 PC; answering a single question took 33 minutes.

Rocky by Rocky
May 26, 2026 - Updated on August 5, 2026
in AI Trends and Related News

When it comes to running AI locally on a computer, most people think you need a large amount of memory, at least an NVIDIA graphics card, or simply buy a Mac Studio. However, the overseas YouTube channel Fully Buffered recently did an interesting test: they took a 20-year-old Intel Pentium 4 platform and challenged it to run an LLM locally, even naming the old machine “NetBurstGPT.” Although it actually managed to run, it was far from practical, with the AI taking 33 minutes just to answer one question.

20 years ago, the Pentium 4 could really run a local LLM: Llama 3.2 was tested successfully, but one question took 33 minutes.

This time, Fully Buffered’s hands-on review covers the Intel Pentium 4 641 processor. It belongs to the late-stage Pentium 4 Cedar Mill core, released in 2006, built on a 65nm process, featuring a 3.2GHz clock speed, 2MB L2 cache, and Hyper-Threading support. This particular unit is the D0 revision, with a TDP of 65W.

By today’s standards, these specs are of course very old, but at the time, within the Pentium 4 family, it was a relatively late and higher-performance variant. Fully Buffered specifically noted that this CPU supports EM64T, which is very important, as that allows it to later run Windows 10 Pro 64-bit and modern AI tools properly:

For memory, the system has four 2GB A-Data PC2-6400 CL5 modules, totaling 8GB of DDR2-800. Additionally, although the computer has a graphics card installed, Fully Buffered stated that he did not want to rely on GPU acceleration and instead wanted to test whether the NetBurst architecture itself could run an LLM. In other words, the AI would run entirely on the Pentium 4 CPU:

Fully Buffered initially tested LM Studio, but the official system requirements clearly state that the Windows x64 version requires a CPU supporting the AVX2 instruction set, and Pentium 4 does not have AVX2. As a result, while LM Studio could open some interfaces and perform certain operations, it would fail when running a model:

He then switched to trying Ollama. Ollama version v0.1.21 supports environments without the AVX instruction set, allowing older CPUs or certain virtualization environments to run normally. For the model, he chose Meta’s Llama 3.2 3B, a relatively lightweight text model:

After the installation was complete, he used Ollama to run the model and entered a simple question: “What’s a Pentium 4?” Then he began waiting. Task Manager also showed 100% running on the CPU, with no GPU being used:

In the end, it actually succeeded; however, this simple question took nearly 33 minutes to answer, with a prompt eval rate of about 0.27 tokens/s and an eval rate of about 0.21 tokens/s:

Fully Buffered also tested Linux Mint to see whether the same Ollama plus Llama 3.2 3B setup would deliver better CPU usage or inference speed in a Linux environment.

As it turned out, no. The speed was actually slower on Linux Mint, with a prompt eval rate of about 0.15 tokens/s and an eval rate of about 0.13 tokens/s:

Finally, he even tried overclocking the Pentium 4, pushing the clock speed from the original 3.2 GHz to 4.3 GHz, with memory speed at around 810 MT/s. After overclocking, it did indeed become faster—by about 20%—with the prompt eval rate increasing to about 0.36 tokens/s and the eval rate reaching about 0.33 tokens/s:

To put these numbers into perspective, Fully Buffered also ran a comparison on his own modern computer using the same Ollama setup and model. His Intel Core i5-12600K was roughly 200 times faster, while an NVIDIA Titan V GPU was roughly 600 times faster.

This comes as no surprise, since the Pentium 4’s NetBurst architecture was designed in the late 1990s.

Full video:

Source: KOCPC Chinese

Tags: aiCPULLMPentium 4processor

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology