NVIDIA and Google’s infrastructure reduces the cost of AI inference

NVIDIA and Google’s infrastructure reduces the cost of AI inference

By Ryan Daws
Publication Date: 2026-04-23 12:19:00

At the Google Cloud Next conference, Google and NVIDIA unveiled their hardware roadmap to address the cost of AI inference at scale.

The companies detailed the new A5X bare metal instances running on NVIDIA Vera Rubin NVL72 rack-scale systems. By co-designing hardware and software, this architecture aims to deliver up to ten times lower inference cost per token compared to previous generations, while achieving ten times higher token throughput per megawatt.

Connecting thousands of processors requires enormous bandwidth to avoid processing delays. The A5X instances address this hardware challenge by combining NVIDIA ConnectX-9 SuperNICs with Google Virgo networking technology.

This configuration scales to 80,000 NVIDIA Ruby GPUs within a single site cluster and up to 960,000 GPUs in a multi-site deployment. Operating at this scale requires sophisticated workload management as routing data across nearly a million parallel processors requires precise…