In recent years, tech giants and research institutions have focused on model training and architectural innovation in AI: from the birth of the Transformer to the rise of multimodal generative models, breakthroughs in AI have primarily occurred at the training and algorithmic levels. Now, however, the spotlight is shifting toward “inference”: how to serve hundreds of millions of users with lower latency and higher efficiency. Google notes that this shift marks the arrival of the “Age of Inference.” With the rise of agentic workflows and surging demand for general-purpose computing, modern AI systems must be able to flexibly coordinate accelerated computing with general compute resources. This trend has prompted Google to rethink chip design, deeply integrating software and hardware to build a new generation of computing platforms purpose-built for inference and training: the Ironwood TPU and the new virtual machine Axion.

Ironwood Arrives: Google’s Seventh-Generation TPU Marks a Technological Leap
Google officially unveiled its seventh-generation TPU in 2025:IronwoodThis chip is regarded as a milestone in the company’s AI hardware development. Ironwood’s design philosophy is to provide both high-performance training and low-latency inference capabilities under extreme loads, meeting the massive computational demands of current large models and generative AI.

According to official data, Ironwood demonstrates remarkable generational performance improvements compared to the previous-generation chip.
-
Compare TPU v5pPeak performance improvement 10 times。
-
Compare TPU v6e (Trillium): in training and inference workloads, performance improvement exceeds 4 times。
This makes Ironwood Google’s most powerful and energy-efficient custom chip to date, capable of supporting a wide range of workloads, from training large language models (LLMs) to real-time generative inference.
Anthropic, Lightricks, Essential AI and other major players are among the first to adopt
Upon its release, Ironwood quickly attracted several leading AI companies to deploy and test it.
Anthropic is one of the earliest partners to adopt Ironwood. James Bradbury, the company’s head of computing, stated: “As our Claude models serve millions of users, our demand for inference performance and training scalability continues to grow. Ironwood’s high cost-performance allows us to scale computing resources more efficiently while maintaining the speed and reliability our users expect.” Anthropic expects to use up to 1 million TPUs to support its model services.
Lightricks The research team used Ironwood’s powerful interchip interconnect network (ICI) to train the multimodal generative model LTX-2, significantly improving efficiency. Company director Yoav HaCohen emphasized: “Ironwood will allow us to produce more realistic images and video content at lower cost, continuing to push the boundaries of open-ended creativity.”
Additionally,Essential AI Philip Monk, the lead of the infrastructure team, also noted: “Ironwood’s efficient scalability and software integration allow us to focus on AI model innovation rather than the bottlenecks of system operations.”
Ironwood’s core: the computing hub of the AI Hypercomputer
Ironwood is not just a single chip, it is Google AI Hypercomputer The core of it all: an integrated supercomputing system that fuses computing, networking, storage, and software.

According to the IDC report, enterprises that adopt AI Hypercomputer can achieve within three years 353% return on investment、Reduce IT costs by 28%, and enhance 55% IT team efficiency。
Ironwood TPU features an ultra-high-density interconnect design in its architecture:
-
Single Superpod Can accommodate 9,216 TPU chips。
-
Transmission bandwidth Gundam 9.6 TB/s。
-
Shared High Bandwidth Memory (HBM) Gundam 1.77 PB。
This massive internet network enables thousands of chips to work in sync, virtually eliminating data transmission bottlenecks. To ensure stability, Google has implemented Optical Circuit Switching (OCS) technology, which can instantly reconfigure the network in the event of an outage, keeping services uninterrupted. Should demand grow, Ironwood can further Jupiter data center network Forming clusters of hundreds of thousands of TPUs, building cloud-scale supercomputer infrastructure.
Ironwood’s hardware performance is further amplified through Google’s software ecosystem. Google has introduced several innovations on the TPU platform:
-
GKE Cluster Director: Provides topology-aware and intelligent scheduling capabilities, enabling TPU clusters to dynamically allocate resources and maintain high elasticity.
-
MaxText open-source frameworkSupports the latest supervised fine-tuning (SFT) and generative reinforcement learning (GRPO) techniques.
-
vLLM on TPU supportAllow developers to seamlessly switch between GPU and TPU, flexibly configuring inference workloads.
-
GKE Inference GatewayOptimize inference latency, reducing time to first token (TTFT) by up to 96% and cutting costs by approximately 30%.
Through these software-layer enhancements, Google has successfully enabled Ironwood to achieve maximum efficiency at every stage of training, fine-tuning, and inference, forming a true “system-level intelligent computing platform.”
Axion Debuts: Arm-Based CPUs Lead the Next Era of General-Purpose Computing
Beyond the AI acceleration led by Ironwood, Google simultaneously launched a Arm Neoverse® architecture of Axion series CPU, redefining the efficiency and flexibility of cloud general-purpose computing.

The two latest released products include:
-
Axion N4A (Preview)Designed for microservices, containerized applications, open-source databases, and data analytics, offering up to … compared to same-class x86 VMs. 2x the value。
-
Axion C4A metal (Preview Version)Google’s first Arm-based bare metal instance, suitable for Android development, automotive systems, or complex simulation environments.
These two products and the existing ones C4A Together they form the complete Axion product portfolio, enabling enterprises to flexibly choose the optimal computing solution based on workload requirements.

Axion’s Real-World Impact: Success Stories from Vimeo, ZoomInfo, and Rise
Vimeo When running video transcoding tests on the Axion N4A, performance improved. 30%It can improve unit economics without requiring changes to the existing architecture.
ZoomInfo After the data processing platform runs on N4A, the cost-performance ratio improves. 60%, significantly accelerating customer data analysis and insight generation.
Rise By migrating to Axion C4A, computing costs decrease. 20%while maintaining low latency and high stability. The company is testing the N4A series to support highly elastic API services and has observed reduced CPU usage. 15%further reducing cloud spending.
Ironwood × Axion: The Perfect Concerto of AI and General-Purpose Computing
Google’s strategy is clear and unambiguous: to Ironwood TPU The “intelligence core” responsible for AI model training and inference, with Axion CPU Handling everyday computing and application-layer tasks, the two together constitute the dual-engine architecture of the AI era.
This vertically integrated design lets businesses achieve both extreme performance and flexibility at the same time. Whether training large language models, deploying generative applications, or handling data analytics and web services, Google Cloud’s compute platform delivers the optimal combination.
Source: KOCPC Chinese