• About Us
King of Computer Media
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us
No Result
View All Result
King of Computer Media
No Result
View All Result

Home - AI Trends and Related News - AI Costs Are Crashing! NVIDIA’s Dual-Track Strategy Pays Off as Era of Sub-$1 Per Million Tokens Arrives

AI Costs Are Crashing! NVIDIA’s Dual-Track Strategy Pays Off as Era of Sub-$1 Per Million Tokens Arrives

KOCPC Editor by KOCPC Editor
April 24, 2026 - Updated on August 5, 2026
in AI Trends and Related News, Latest Technology News

As the global AI industry’s focus shifts from model training to inference applications, the traditional AI infrastructure market, which has been priced based on GPU hourly rental rates, is rapidly transitioning toward a new model centered on cost per million tokens. NVIDIA, at the end of 2025, approximately $20 billion (approximately NT$650 billion) at a staggering price, acquired AI inference chip startup Groq. This is not only NVIDIA’s largest acquisition in history, but now appears to be a well-considered strategic move: by integrating Groq’s LPU technology, NVIDIA has further cemented its seemingly unassailable market dominance in the inference chip sector.

The Pricing Model Revolution: From GPU Hourly Rental to Token-Based Pricing

Traditionally, the pricing of AI computing infrastructure depends on the GPU model and whether capacity is reserved or on-demand. According to data provided by experts from cloud infrastructure provider Nebius in an interview with AlphaSense, the current on-demand pricing for various generations of NVIDIA GPUs is as follows:

  • H100per hour $2.95(approx. NT$96)
  • H200per hour $3.50(approx. NT$114)
  • Blackwell B200per hour $4.90 to $6.50(approximately NT$159 to NT$211)

If signing a 1-2 year long-term contract with a guaranteed purchase of at least 10,000 GPUs, the price will drop significantly: H100 to $1.50H200 dropped to $2.20 USDB200 drops to $3.50That’s all.

However, experts point out that the market is rapidly shifting to pricing models based on “per million tokens.” The fundamental driving force behind this shift is:Inference now accounts for 90% to 95% of everyday AI workloadsAs various models become increasingly mature, enterprises heavily rely on pre-trained models or API services rather than developing their own models from scratch. The inference efficiency of large language models has replaced training performance as the key factor determining AI costs.

The Strategic Significance of a $20 Billion Acquisition of Groq

In December 2025, NVIDIA’s approximately $20 billion acquisition of Groq’s chip assets and engineering team shocked the industry. This deal carries two layers of core significance:

  1. Strategic Supplement to the LPU ArchitectureGroq’s LPU (Language Processing Unit) employs a unique “Software-Defined Architecture,” completely different from NVIDIA’s traditional GPU approach. What NVIDIA sees is Groq’s unmatched advantage in inference latency.
  2. Dual-track reasoning strategyAfter the acquisition, NVIDIA established a dual-track product line combining “Blackwell high-throughput training + Groq LPU ultra-low-latency inference.” At the 2026 GTC conference, NVIDIA Vice President Ian Buck openly stated LPU’s inference efficiency is more than 5 times that of the H100。

NVIDIA Internal Dual-Architecture Comparison

Indicator NVIDIA Blackwell B200 NVIDIA Groq LPU (integrated after acquisition)
Cost per million tokens $0.25 USD(approx. NT$8.1) $0.05–$0.10 USD(about NT$1.6–3.3)
Token Output Speed 450 tokens/sec 800 tokens/sec
Main application scenarios Training + Enterprise-grade Reasoning High-speed inference (ultra-low latency)

This means that NVIDIA now simultaneously has Two completely different inference chip architecturesBlackwell handles high-intensity training and general-purpose workloads, while Groq LPU focuses on real-time inference scenarios requiring ultra-low latency. This isn’t “challenger vs. incumbent”—it’s NVIDIA’s complete domination of the entire AI infrastructure.

Blackwell’s 35x Generational Leap

Meanwhile, NVIDIA’s own Blackwell architecture has also demonstrated a staggering generational leap. According to official benchmarks, the cost per million tokens has dropped from the Hopper generation (H200) to Blackwell, from $4.20 USDdropped sharply to $0.12, a decline of up to 35 times。

More compelling is the actual throughput comparison: when running the DeepSeek-R1 model, the Blackwell GB300 NVL72 configuration can generate per GPU 6,000 TokensThe H200 only manages 90 tokens. Although Blackwell costs approximately $2.65 per GPU hour (higher than the H200’s $1.41), the actual token output differs by nearly 67 times.

At the same time, Blackwell’s energy efficiency performance is equally impressive: per megawatt can generate 2.8 million TokensIt’s Hopper’s more than 50 timesFor enterprise data centers with long-term power contracts, this advantage is far more decisive than chip pricing.

Jensen Huang’s “Token Economics” and “AI Factories”

At GTC 2026, NVIDIA CEO Jensen Huang devoted a significant portion of his keynote to promoting the new concept of “Token Economics.” He argued that data centers have evolved into “AI Token factories,” and that the standard for measurement should shift away from FLOPS or GPU hours toward “cost per Token.”

Jensen Huang further outlined a bold vision for the future: “In the future, every engineer, in addition to their salary, will also have a Token budgetI will also allocate half of my salary for them to use, enabling them to achieve 10x productivity through AI.

He emphasized that NVIDIA’s “Inference Iceberg” framework—including FP4 precision support, Speculative Decoding, KV-Cache offloading, and the Dynamo service layer that has just entered mass production—is what truly determines actual token throughput. Chips lacking these software-layer optimizations, no matter how high their peak specifications, will have significantly reduced actual token output.

Rubin Platform: The Next Milestone

NVIDIA’s layout extends beyond Blackwell. Expected to end of 2026the next generation launched Rubin PlatformAccording to analysis, it claims to be able to reduce token costs by an additional 90% on the basis of Blackwell. This shows that NVIDIA is continuing to widen the gap with competitors through its approach of developing multi-generational architectures simultaneously.

Enterprise Choices: Build vs. Rent vs. Embrace AI Factory

For enterprises, the good news is that major cloud partners—including CoreWeave, Nebius, Nscale, and Together AI—have already deployed Blackwell infrastructure at scale. Enterprises can enjoy advanced inference economics at less than $1 per million tokens without massive capital investment. Gartner further predicts that AI token costs will drop by more than 90% by 2030, and NVIDIA’s data shows this price reduction trend is already accelerating.

Conclusion: NVIDIA’s Complete Victory

From hourly GPU rentals to token-based pricing, from H100 to Blackwell to Rubin, and from GPU to LPU dual-architecture deployment: NVIDIA is building an all-encompassing AI infrastructure empire covering training, inference, and latency-sensitive scenarios. The $20 billion acquisition of Groq is just one move in this grand strategy.

In the future, “What is the cost per million tokens?” will replace “Which GPU should we use?” as the primary question for enterprise decision-making. And in this new era, NVIDIA already holds all the answers.

Source, KOCPC Chinese

Tags: BlackwellgroqNVIDIAVera Rubin黃仁勳

Recent Posts

  • The Xiaomi Pad 8S Pro has passed network access certification and will debut with the self-developed XRING O3 chip.
  • The entire Google Pixel 11 lineup has been leaked! Official promotional renders of the Pixel 11 Pro XL have also surfaced
  • Are Chinese phone battery capacities falsely labeled? A brief look at the “capacity locking” phenomenon in Chinese silicon-carbon batteries.
  • NCC is leaderless, recklessly sending out national-level alert messages!?
  • What does “QR” in QR Code mean?

Recent Comments

No comments to show.
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology

No Result
View All Result
  • Home
  • Tech News
  • AI News
  • Apps & Tutorials
  • Mobile & Telecom
  • Lifestyle
  • About Us

We welcome partnership inquiries and product review opportunities from smartphone manufacturers, iPhone accessory brands, and app developers.koc kocpc.com.tw|Privacy Policy |Hosting & Maintenance: Fast Line Taiwan, A-Chang Digital Technology