By Scott Guthrie
Publication Date: 2026-01-26 16:00:00
Today we are proud to introduce Maia 200, a groundbreaking inference accelerator designed to significantly improve the economics of AI token generation. Maia 200 is an AI inference powerhouse: an accelerator based on TSMC’s 3nm process with native FP8/FP4 tensor cores, a redesigned storage system with 216GB HBM3e at 7TB/s and 272MB on-chip SRAM, and data movement engines that keep large models fast and highly utilized. This makes Maia 200 the most powerful first-party silicon of any hyperscaler, with three times the FP4 performance of the third-generation Amazon Trainium and the FP8 performance over Google’s seventh-generation TPU. Maia 200 is also the most efficient inference system Microsoft has ever deployed, with 30% better performance per dollar than the latest generation hardware in our fleet today.
Maia 200 is part of our heterogeneous…



