AI’s Cloud Cost Reckoning: How Vendors Are Trying To Tame Token, GPU and Datacenter Bills — Virtualization Review

AI’s Cloud Cost Reckoning: How Vendors Are Trying To Tame Token, GPU and Datacenter Bills — Virtualization Review

By By David Ramel05/29/2026
Publication Date: 2026-05-29 00:00:00

In-Depth

AI’s Cloud Cost Reckoning: How Vendors Are Trying To Tame Token, GPU and Datacenter Bills

Cloud providers are facing a two-sided AI cost problem: They have to keep building the infrastructure required to support rising demand, while giving enterprise customers enough controls to keep AI workloads financially manageable.

The result is a new phase in cloud AI, one in which providers are not only competing on model access and performance, but also on cost controls. Prompt caching, context caching, model routing, provisioned throughput, reserved capacity, service tiers, batch processing and custom AI chips are becoming part of the cloud AI product stack.


Prompt Caching
[Click on image for larger view.] Prompt Caching(source: AWS).

The shift is happening as AI demand continues to drive large infrastructure investments. In its fiscal 2026 third-quarter earnings release, Microsoft reported $82.9 billion in revenue, said Microsoft Cloud revenue reached $54.5 billion and said its AI business surpassed a $37 billion annual revenue run rate. The same release showed additions to property and equipment of $30.876 billion for the quarter and $80.146 billion for the first nine months of the fiscal year.

Those figures illustrate the scale of the cloud AI buildout. AI…