Google Cloud Splits AI TPUs to Challenge Nvidia | eWEEK

Google Cloud Splits AI TPUs to Challenge Nvidia | eWEEK

By eWEEK Staff
Publication Date: 2026-04-23 10:19:00

eWeek content and product recommendations are
editorially independent. We may make money when you click on links
to our partners.
Learn More

Google Cloud is competing more directly with Nvidia in AI infrastructure.

At Cloud Next 2026, Google introduced TPU 8t for training and TPU 8i for inference, splitting its latest AI chips between building large models and serving them in production.

Why Google is splitting training and inference

According to Google Cloud’s eighth-generation TPU overview, TPU 8t is built for training, and TPU 8i is built for inference, with both running on Google’s Axion Arm-based processors. Google says the split reflects a growing divide between the hardware needed to train large models and the hardware needed to run AI services efficiently once those models are in use.

Training large models depends on scale, dense compute, and fast communication across large clusters. Inference places greater emphasis on latency, memory use, power efficiency, and the cost of serving requests over time.

In its technical deep dive on TPU 8t and TPU 8i, Google says TPU 8i delivers up to 80% better performance per dollar and twice the performance per watt for inference than the prior generation. Those figures come from Google rather than independent testing, so they are best read as product claims until customers can measure them in live deployments.

Google is also packaging the chips inside a broader system. Its AI Hypercomputer