• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - Google Launches Gemma 4 MTP Drafters, Boosting Local AI Inference Speed by 3x

Google Launches Gemma 4 MTP Drafters, Boosting Local AI Inference Speed by 3x

KOCPC Editor by KOCPC Editor
May 8, 2026 - Updated on August 5, 2026
in AI Trends and Related News, Latest Technology News

On May 5, Google announced the launch of Multi-Token Prediction (MTP) Drafters technology for its open-source model Gemma 4. This technology leverages Speculative Decoding to deliver up to 3x faster inference speed when Gemma 4 runs locally, without any compromise in output quality.

What is MTP and speculative decoding?

Traditional large language models (LLMs) use a reasoning approach called “autoregressive generation,” which generates only one token at a time, and each token requires moving billions of parameters from VRAM to the compute unit, causing the compute unit to spend most of its time idle and waiting. This is the fundamental reason why local AI responses are slow—not because the model isn’t smart enough, but because memory bandwidth has become the bottleneck.

The core concept of MTP is “speculative decoding,” a technique first proposed by Google researchers in a 2022 paper titled “Fast Inference from Transformers via Speculative Decoding.” The approach pairs the main model with a lightweight “draft model” (Drafter), which quickly predicts multiple possible token sequences, and the main model then validates these predictions in parallel in a single pass.

If the main model agrees with the draft’s predictions, it accepts the entire sequence and completes it in a single forward pass, while also generating one additional token of its own. This means the application can output an entire draft sequence plus one new token in the time it would normally take to produce just a single token.

According to Google’s official documentation, the Drafter model directly uses the main model’s Key-Value Cache, so there’s no need to recalculate the context information already processed by the main model, which significantly reduces the additional overhead.

How much actual performance improvement is there?

According to Google’s published benchmark data, the speedup effects brought by MTP on different hardware are as follows:

On Pixel phones, the lightweight E2B model achieves a 2.8x speedup, while the E4B model reaches 3.1x. On Apple M4 chips, the Gemma 4 31B model gains a 2.5x speedup. On mid-to-high-end NVIDIA RTX PRO 6000 graphics cards, the 26B MoE model with MTP delivers approximately twice the original output speed.

 Your browser does not support video playback.

Note that 3x is the theoretical upper limit, and the actual speedup depends on hardware specifications and usage scenarios. For example, on Apple Silicon, batch processing 4 to 8 requests can achieve approximately 2.2x speedup. Google is also optimizing for different hardware architectures, including Apple Silicon and NVIDIA A100.

The significance for developers and users

For developers, inference speed is often the biggest bottleneck when deploying AI models to production environments. The introduction of MTP brings practical improvements to the following scenarios:

Instant messaging appIn the past, local AI conversations always suffered from noticeable lag, but now they can operate at near-real-time speeds.
Code AssistantThe wait time for line-by-line reasoning during development has been significantly reduced, making offline code assistance truly usable.
Voice interactionVoice interfaces are extremely latency-sensitive, and MTP can bring local voice assistants’ response times up to a practical level.
Agentic tasks(Agentic Workflows): AI agents requiring multi-step planning will see a significant reduction in wait time between each step.

Additionally, since the main model still handles final verification, MTP does not affect output quality in any way—what Google refers to as “Zero Quality Degradation.” The correctness and reasoning capabilities of the model are entirely ensured by the main model, while Drafter is only responsible for acceleration.

How to get and use

MTP Drafters and Gemma 4 alikeLicensed under Apache 2.0As of today, model weights can be downloaded from platforms such as Hugging Face, Kaggle, and Google AI Edge Gallery. Supported frameworks include Hugging Face Transformers, MLX (Apple Silicon), vLLM, SGLang, and Ollama.

For most users, the simplest way to try it is via Ollama:ollama run gemma4:31b-coding-mtp-bf16Android and iOS users can also experience it through the Google AI Edge Gallery app.

Source: KOCPC Chinese

Tags: Gemma 4GoogleMTPMTP DraftersSpeculative decoding

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology