PARTNER CONTENT Google is the first of the hyperscalers, and many would argue the most innovative across the many businesses and workloads that have driven its collective for the past three decades. The pressure on the Chocolate Factory to reduce costs and to beat Moore's Law no matter the limits of physics is unrelenting. That's good news for Google Cloud customers because they reap the innovations the wider Google business creates for its own sake.
Following the debut of Arm delivered its Neoverse IP blocks in 2018, the architecture matured in tandem with a rapidly expanding software ecosystem. Driven by the deployment of advanced Arm server CPUs, the industry shifted toward full ecosystem parity—equipping modern operating systems, developer tools, and critical enterprise applications with first-class, native support for Arm. That is when Google jumped in with its homegrown Axion Arm server CPUs two years ago, driving down server CPU costs and tailoring performance for specific workloads inside Google.
The reason was simple. Like every other enterprise on the planet, the search giant needed to modernize its server fleet so it would be more efficient and less costly, freeing capital to reinvest in AI. All the benefits of Google's Axion server CPUs and their related Titanium data processing units (DPUs) accrue to any customer who wants virtual or bare metal compute for any workload they can think of, and all they have to do is buy capacity on Google Cloud to get the most tailored compute Google itself uses.
"Google is investing in custom silicon across our datacenter offerings, both for internal use as well as customer exposed use," says Mo Farhat, group product manager at Google Cloud who leads the Axion product and go to market strategy.
Farhat describes the Axion CPUs as one of the important phases of this innovation. The insight it gets from running countless workloads and the vertical integration from silicon to service are huge advantages.
The Performance And Efficiency To Meet The NeedsApril 2024 saw Google Cloud reveal first Axion processor, powering its C4A instances. Google chose the "Demeter" Arm Neoverse V2 core as the basis of its design.
Demeter's big draw was its substantial vector math performance and capable Arm integer core based on the Armv9.0-A architecture. It has 64 KB of L1 instruction cache, 64 KB of L1 data cache, and a 2 MB private L2 cache per core. There is also an 80 MB L3 cache shared across the 72 cores on the CPU.
The Demeter cores do not implement simultaneous multithreading (SMT). That helps with security by limiting the attack surface and offering absolute isolation at the core. It also delivers more deterministic performance per thread than cores that do implement SMT.
The Demeter V2 cores have four 128-bit SVE2 vector units, which are, ironically enough, easier to keep busy than the two 256-bit vector units in the "Zeus" V1 core. The SVE2 unit also supports MatMul (I8MM) operations and pointer authentication (PAC).
The C4A instance are aimed at in-memory caches, databases, high-traffic web and application servers, CPU-based high performance computing, and other workloads where core counts and complex math are de rigueur, not du jour. The C4A variant supports up to 576 GB of DDR5-5600 main memory against its 72 cores, while the C4A Metal version goes to 768 GB of DDR5-5600 memory across 96 cores.
Ethernet networking out of these instances is 100 Gb/sec, and for the C4A instance, there is up to 6 TB of local flash that can drive up to 10.4 GB/sec of read throughput and 2.4 million random read I/O operations per second. The C4A instances became available in October 2024.
Google expanded the Axion processor family with the N4A instances that were previewed last fall and became available in January of this year. They support up to 512 GB of DDR5 memory .
The N4A instances support a new Google Cloud technology called Custom Machine Types, which means the vCPU count and memory capacity can be configured and priced independently so customers can tailor instances for their application needs without leaving any vCPUs or memory capacity stranded.
The N4A instances have 50 Gb/sec of Ethernet connectivity out of the Titanium DPU . The N4A instances are aimed at web servers, GKE containerized applications, CI/CD application development, and lightweight data analytics workloads.
The typical Axion processor is between 3.2 GHz and 3.4 GHz. It's set and locked by the Titanium DPU, which handles network virtualization, storage virtualization (for both the external Hyperdisk block storage service and the local Titanium SSD flash in some Axion instance nodes), and hardware-based root of trust security through the Titan security chips on the DPU.
A thin hypervisor layer still runs on the C4A and N4A instances, but the C4A Metal instance is bare metal without a hypervisor. By offloading these functions to the Titanium DPU, more cores are available on the Axion chip to do actual work. In the old days, before DPUs, you could lose 30 percent to 40 percent of the cores on a chip to these network and storage functions.
Performance and price comparisons are where the Axion processors get tested against the rest of the market.
Google Cloud says the Arm-based C4A instances have up to 65 percent better bang for the buck and 60 percent better performance per watt compared to x86 based C4 instances.
For the Arm-based N4A instances, Google Cloud says that on the SPECrate2017 integer test, Axion CPU delivers 105 percent better price/performance than comparable x86-based N4 instances, and up to 90 percent better price/performance for web servers using the Nginx reverse proxy benchmark as the touchstone.
For Java workloads as gauged by the SPECjbb2015 benchmark, value for dollar is 85 percent better than for the x86-based N4 instances. As gauged by the MySQL throughput benchmark, the N4A instances delivered 20 percent better bang for the buck than x86-based N4 instances.
Google benefits from using pre-made Arm Neoverse cores and perhaps other Arm IP rather than creating its own custom cores, which it is perfectly technically capable of doing. Because Google can use the entire Arm software stack, it can extend its build system to port its own applications from x86 platforms to Arm based Axion CPUs.
With over 100,000 applications to move, this is a truly Herculean task. So far, 30,000 Google applications have been ported to Axion processors by last October, starting with the biggies — the Spanner global file system, the BigTable database layer that runs atop it, and the BigQuery data warehouse, plus YouTube video streaming and Gmail.
"We are in a world where more and more of the software ecosystem is being developed around Arm, more and more of the hardware ecosystem is being developed around Arm," Dermot O'Driscoll, Vice President of Go to Market and Customer Solutions at the Physical AI, at Arm, tells The Next Platform.
"When you have the four US major clouds and hyperscalers dedicated to Arm architecture, you are going to see a lot more developers wanting to be working on - Arm platforms rather than on a legacy architecture with a legacy cost structure," he continues. "The custom SoCs that Google has designed are created to do what the company thinks they need to do well, and Google doesn't have to carry architectural baggage that is not interesting to their own needs and that of their customers. So that level of fine tuning the hardware to meet their software needs, it's unique."
Sponsored by Arm.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | With TPU 8, Google Makes GenAI Systems Much Better, Not Just Bigger | 0 | 18.33 | 24-04-2026 |
| 2 | A year in, Google wants its Axion processors to feel like a scheduling decision | 0 | 12.38 | 15-04-2026 |
| 3 | NextSilicon Takes Aim At CPUs And GPUs With “Maverick-2” Dataflow Engine | 0 | 10 | 22-10-2025 |
| 4 | Broadcom Helps CPU And XPU Makers Go Vertical With Compute | 0 | 10 | 05-05-2026 |
| 5 | AMD Catches The Agentic AI Wave And Will Ride It Up Masterfully | 0 | 10 | 05-08-2026 |
| 6 | AI Hosts And Sandboxes Save Intel’s Datacenter CPU Cookies | 0 | 10 | 28-07-2026 |
| 7 | Arm and Google offer a smarter option to run agentic AI workloads | 0 | 14.95 | 17-07-2026 |
| 8 | With Taalas, AMD Can Bake AI Inference Directly Into Its Chippery | 0 | 10 | 07-08-2026 |
| 9 | Of Course Meta Platforms Is Going To Be A Cloud | 0 | 10 | 01-07-2026 |
| 10 | Nvidia Accelerates Chip Engineering With AI Agents | 0 | 10 | 27-07-2026 |