By Timothy Prickett Morgan
Publication Date: 2026-04-24 16:56:00
Here
is how you know that GenAI training and GenAI inference are very different
computing and networking beasts, and diverging more with each passing day:
Google has just forked its Tensor Processing Unit, or TPU, designs for these
two workloads, the very first time in more than a decade that TPU systems of
the same generation were truly architecturally distinct from each other.
To
be fair, Google has had versions of prior TPUs, specialized and homegrown AI
accelerators that the search engine, advertising, media mogul, and now AI model
maker first put into the field way back in 2015. Google got into the TPU racket
because there was no way machine learning algorithms could be added to its
applications using CPUs or GPUs without doubling its datacenter footprint. But a
lot of the time these different TPUs were just bin sorts on a designs that were
primarily aimed at AI training, with the lower bin parts being used for
inference. All of this machine learning was much…




