• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - Latest Technology News - No need to buy high-end large-capacity VRAM graphics cards, Intel’s new driver can use up to 93% of system memory to run large models

No need to buy high-end large-capacity VRAM graphics cards, Intel’s new driver can use up to 93% of system memory to run large models

KOCPC Editor by KOCPC Editor
May 4, 2026 - Updated on August 5, 2026
in Latest Technology News

For users who want to run AI large language models (LLM) locally, insufficient display memory (VRAM) capacity has always been the biggest pain point. In the past, if you wanted to smoothly run AI models with billions of parameter levels, you often had to spend a lot of money to buy high-end professional graphics cards equipped with large-capacity VRAM. However, the latest driver update released by Intel will completely change the rules of this game.

In the latest Arc series driver, Intel has introduced a new feature called “Shared GPU Memory Override” function, allowing the Intel display chip to call the highest 93% The system memory (RAM) is used as display memory. This means that on a computer with 64GB RAM, the GPU can get up to about 59.5GB of available space; on a high-end platform with 128GB RAM, more than 112GB of VRAM capacity can be obtained.

Break 50% of hardware barriers

Under traditional architecture, Intel laptops usually divide the system memory into two: 50% is allocated to the operating system and applications, and the other 50% is reserved for the GPU as shared display memory. This invisible boundary has caused many users to frequently run into obstacles when running large AI models.

In fact, AMD has already used its Variable Graphics Memory (VGM) technology to allow Ryzen platform users to manually adjust the VRAM allocation ratio through the Adrenalin driver or BIOS. For example, the Ryzen AI Max (Strix Halo) platform can allocate up to 96GB to the iGPU in a 128GB RAM system. Bob Duffy, Intel’s AI Playground application director, has previously shared this update on the X platform, emphasizing that this is a major benefit for users who run AI chatbots and image generation applications locally.

If you have Intel Core Ultra and are doing AI, you’re going to want to update to latest Intel Arc driver… because this pic.twitter.com/4BlTqW1RCo

— bobduffy 🖱️🎨 (@bobduffy) August 14, 2025

This function originally existed, but the new driver version directly reaches the 93% upper limit.

This feature comes from the latest version of Intel’s driver (version number 32.0.101.8517), which is mainly targeted at systems equipped with integrated Arc Pro GPU. Prior to this, Intel had already raised the allocation limit for Core Ultra Series 2 processors to 87% last year, and the latest version has further pushed it to 93%.

The hardware currently supported by this update includes:

  • Intel Arc Pro B390 and Arc Pro B370 built-in GPU
  • Intel Arc Pro A-series and B-series discrete graphics cards
  • Intel Core Ultra Meteor Lake, Lunar Lake, Arrow Lake series processors
  • Original Intel Arc A series discrete graphics cards will also benefit simultaneously.

In terms of software, it fully supports multiple major versions of Windows 10 and Windows 11.

Practical combat: Which AI models can run now?

The biggest highlight of this update is that it allows models that could only be run on professional workstations to now be run smoothly on general commercial laptops:

32GB RAM system: Can run Qwen series 32B models (4-bit quantization) and maintain a comfortable context window.

64GB RAM system: Can run heavyweight models such as Llama 3 70B, and still has enough space for KV cache and system stability.

128GB RAM high-end platform: More than 112GB of available space, enough to handle most consumer-level AI application scenarios.

For AI models, the more VRAM, the chatbot with larger parameters can be run, more in-depth answers can be produced, and a larger number of input and output tokens can be processed.

Speed ​​Bottleneck: Bandwidth Still Key

Although the capacity issue has been resolved, the speed aspect still needs to be viewed with caution. There is still a gap between the transfer speed of system memory (DDR5/LPDDR5X) and dedicated display memory (GDDR6/HBM).

Currently, the Intel Core Ultra Series 3 (Panther Lake) processor supports LPDDR5X-9600 memory, which can provide a bandwidth of approximately 150 GB/s. In comparison, AMD’s Strix Halo can reach 256 GB/s of bandwidth with its 256-bit memory bus, while Apple Silicon M5 Max’s Unified Memory Architecture (UMA) can reach 614 GB/s in one fell swoop.

The biggest advantage of Apple’s UMA architecture is that it breaks away from the traditional memory partition restrictions in the x86 world. The entire memory pool can be accessed natively by the CPU and GPU without setting hard partition boundaries. This keeps Apple Silicon at the forefront of AI inference performance.

However, for personal AI applications that emphasize “localization and privacy”, Intel’s update is enough to fill the huge performance gap. After all, for most users, “running fast” is more important than “running fast”. Being able to use a 64GB RAM laptop to run large models locally instead of sending data to the cloud for processing is a big improvement in itself.

Price is the real killer

Currently, the prices of professional graphics cards equipped with large-capacity VRAM on the market are still high. An NVIDIA RTX 6000 Ada Generation (48GB VRAM) costs over NT$150,000, while a 64GB DDR5 memory module costs just a few thousand dollars. Intel’s solution allows users to spend less money and obtain greater AI model carrying capacity.

Although the bandwidth of system memory is not as good as that of dedicated VRAM, for large parametric models that need to be resident in the background, this “space-for-feasibility” strategy has indeed created a more cost-effective path for the AI ​​PC era. In the future, as DDR5 memory bandwidth continues to increase and Intel’s subsequent drivers continue to be optimized, this type of system memory acting as display memory may become a mainstream configuration for AI computing.

Source: KOCPC Chinese

Tags: Arc Pro B370Arc Pro B390INTELShared GPU Memory Override

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology