Maia 200: The AI ​​Accelerator Built for Inference – The Official Microsoft Blog

Maia 200: The AI ​​Accelerator Built for Inference – The Official Microsoft Blog

By Scott Guthrie
Publication Date: 2026-01-26 16:00:00

Microsoft’s next-generation AI accelerator gives Azure an advantage in running AI models faster and more cost-effectively.

Today we are proud to introduce Maia 200, a groundbreaking inference accelerator designed to significantly improve the economics of AI token generation. Maia 200 is an AI inference powerhouse: an accelerator based on TSMC’s 3nm process with native FP8/FP4 tensor cores, a redesigned storage system with 216GB HBM3e at 7TB/s and 272MB on-chip SRAM, and data movement engines that keep large models fast and highly utilized. This makes Maia 200 the most powerful first-party silicon of any hyperscaler, with three times the FP4 performance of the third-generation Amazon Trainium and the FP8 performance over Google’s seventh-generation TPU. Maia 200 is also the most efficient inference system Microsoft has ever deployed, with 30% better performance per dollar than the latest generation hardware in our fleet today.

Maia 200 is part of our heterogeneous…