After tokenmaxxing backlash, Microsoft moves more agent work onto the desk

After tokenmaxxing backlash, Microsoft moves more agent work onto the desk

By Brian Buntz
Publication Date: 2026-10-07 23:46:00

Microsoft’s local MAI-Code-1.1-Flash generates roughly 60 tokens per second with short prompts on Surface Laptop Ultra, versus just under 40 at a 256K-token prompt. The figures come from Microsoft’s testing on a synthetic code-generation workload. Image: Microsoft

2026 saw the rise, and quick fall, of tokenmaxxing, the practice of enterprise companies treating AI spend volume as a yardstick for productivity. This summer, Microsoft put its divisions on AI token budgets and told engineers that “tokenmaxxing is not what we are optimizing for.”

Now, in its Oct. 7 announcements, it offered customers a way to avoid the same sticker shock by moving more of the agents’ work onto the PC. It imagines always-on mini PCs running agents, and shared Nvidia DGX Stations, desk-side AI workstations, serving 32 or more people at once. The cloud stays available for work that needs it.

Coding is an early test of Microsoft’s proposition. The company is bringing its MAI-Code-1.1-Flash model…