As the global AI industry’s focus shifts from model training to inference applications, the traditional AI infrastructure market, which has been priced based on GPU hourly rental rates, is rapidly transitioning toward a new model centered on cost per million tokens. NVIDIA, at the end of 2025, approximately $20 billion (approximately NT$650 billion) at a staggering price, acquired AI inference chip startup Groq. This is not only NVIDIA’s largest acquisition in history, but now appears to be a well-considered strategic move: by integrating Groq’s LPU technology, NVIDIA has further cemented its seemingly unassailable market dominance in the inference chip sector.

The Pricing Model Revolution: From GPU Hourly Rental to Token-Based Pricing
Traditionally, the pricing of AI computing infrastructure depends on the GPU model and whether capacity is reserved or on-demand. According to data provided by experts from cloud infrastructure provider Nebius in an interview with AlphaSense, the current on-demand pricing for various generations of NVIDIA GPUs is as follows:
- H100per hour $2.95(approx. NT$96)
- H200per hour $3.50(approx. NT$114)
- Blackwell B200per hour $4.90 to $6.50(approximately NT$159 to NT$211)
If signing a 1-2 year long-term contract with a guaranteed purchase of at least 10,000 GPUs, the price will drop significantly: H100 to $1.50H200 dropped to $2.20 USDB200 drops to $3.50That’s all.
However, experts point out that the market is rapidly shifting to pricing models based on “per million tokens.” The fundamental driving force behind this shift is:Inference now accounts for 90% to 95% of everyday AI workloadsAs various models become increasingly mature, enterprises heavily rely on pre-trained models or API services rather than developing their own models from scratch. The inference efficiency of large language models has replaced training performance as the key factor determining AI costs.
The Strategic Significance of a $20 Billion Acquisition of Groq
In December 2025, NVIDIA’s approximately $20 billion acquisition of Groq’s chip assets and engineering team shocked the industry. This deal carries two layers of core significance:
- Strategic Supplement to the LPU ArchitectureGroq’s LPU (Language Processing Unit) employs a unique “Software-Defined Architecture,” completely different from NVIDIA’s traditional GPU approach. What NVIDIA sees is Groq’s unmatched advantage in inference latency.
- Dual-track reasoning strategyAfter the acquisition, NVIDIA established a dual-track product line combining “Blackwell high-throughput training + Groq LPU ultra-low-latency inference.” At the 2026 GTC conference, NVIDIA Vice President Ian Buck openly stated LPU’s inference efficiency is more than 5 times that of the H100。
NVIDIA Internal Dual-Architecture Comparison
| Indicator | NVIDIA Blackwell B200 | NVIDIA Groq LPU (integrated after acquisition) |
|---|---|---|
| Cost per million tokens | $0.25 USD(approx. NT$8.1) | $0.05–$0.10 USD(about NT$1.6–3.3) |
| Token Output Speed | 450 tokens/sec | 800 tokens/sec |
| Main application scenarios | Training + Enterprise-grade Reasoning | High-speed inference (ultra-low latency) |
This means that NVIDIA now simultaneously has Two completely different inference chip architecturesBlackwell handles high-intensity training and general-purpose workloads, while Groq LPU focuses on real-time inference scenarios requiring ultra-low latency. This isn’t “challenger vs. incumbent”—it’s NVIDIA’s complete domination of the entire AI infrastructure.

Blackwell’s 35x Generational Leap
Meanwhile, NVIDIA’s own Blackwell architecture has also demonstrated a staggering generational leap. According to official benchmarks, the cost per million tokens has dropped from the Hopper generation (H200) to Blackwell, from $4.20 USDdropped sharply to $0.12, a decline of up to 35 times。
More compelling is the actual throughput comparison: when running the DeepSeek-R1 model, the Blackwell GB300 NVL72 configuration can generate per GPU 6,000 TokensThe H200 only manages 90 tokens. Although Blackwell costs approximately $2.65 per GPU hour (higher than the H200’s $1.41), the actual token output differs by nearly 67 times.

At the same time, Blackwell’s energy efficiency performance is equally impressive: per megawatt can generate 2.8 million TokensIt’s Hopper’s more than 50 timesFor enterprise data centers with long-term power contracts, this advantage is far more decisive than chip pricing.
Jensen Huang’s “Token Economics” and “AI Factories”
At GTC 2026, NVIDIA CEO Jensen Huang devoted a significant portion of his keynote to promoting the new concept of “Token Economics.” He argued that data centers have evolved into “AI Token factories,” and that the standard for measurement should shift away from FLOPS or GPU hours toward “cost per Token.”
Jensen Huang further outlined a bold vision for the future: “In the future, every engineer, in addition to their salary, will also have a Token budgetI will also allocate half of my salary for them to use, enabling them to achieve 10x productivity through AI.
He emphasized that NVIDIA’s “Inference Iceberg” framework—including FP4 precision support, Speculative Decoding, KV-Cache offloading, and the Dynamo service layer that has just entered mass production—is what truly determines actual token throughput. Chips lacking these software-layer optimizations, no matter how high their peak specifications, will have significantly reduced actual token output.
Rubin Platform: The Next Milestone
NVIDIA’s layout extends beyond Blackwell. Expected to end of 2026the next generation launched Rubin PlatformAccording to analysis, it claims to be able to reduce token costs by an additional 90% on the basis of Blackwell. This shows that NVIDIA is continuing to widen the gap with competitors through its approach of developing multi-generational architectures simultaneously.
Enterprise Choices: Build vs. Rent vs. Embrace AI Factory
For enterprises, the good news is that major cloud partners—including CoreWeave, Nebius, Nscale, and Together AI—have already deployed Blackwell infrastructure at scale. Enterprises can enjoy advanced inference economics at less than $1 per million tokens without massive capital investment. Gartner further predicts that AI token costs will drop by more than 90% by 2030, and NVIDIA’s data shows this price reduction trend is already accelerating.
Conclusion: NVIDIA’s Complete Victory
From hourly GPU rentals to token-based pricing, from H100 to Blackwell to Rubin, and from GPU to LPU dual-architecture deployment: NVIDIA is building an all-encompassing AI infrastructure empire covering training, inference, and latency-sensitive scenarios. The $20 billion acquisition of Groq is just one move in this grand strategy.
In the future, “What is the cost per million tokens?” will replace “Which GPU should we use?” as the primary question for enterprise decision-making. And in this new era, NVIDIA already holds all the answers.