Microsoft seeks capacity for 300,000 Maia 300 chips to cut AI costs

Microsoft is seeking TSMC capacity for more than 300,000 chips in 2027 after deploying Maia 200 in only two U.S. data centers.

By · Published

Why it matters

Maia 300 could lower Microsoft's inference costs and reserve Nvidia systems for paying Azure customers, but the strategy needs manufacturing scale and an outside customer.

An intricate AI chip architecture, monumental in scale and complexity (Woodblock print in the manner of mid-century propaganda posters, with flat planes of color and bold silhouettes)

Microsoft plans to unveil its Maia 300 AI accelerator this fall and is discussing manufacturing capacity for more than 300,000 chips in 2027, a sharp increase meant to turn CEO Satya Nadella's custom-silicon program into a credible alternative to Nvidia.

The Information reported Monday that Microsoft could introduce Maia 300 as early as September. Microsoft has produced only tens of thousands of the current Maia 200 chip, according to the report, and was still running those accelerators in two U.S. data centers as of late July.

Microsoft has been negotiating with Taiwan Semiconductor Manufacturing Co. for capacity to make more than 300,000 Maia 300 chips next year, according to The Information, citing a person with direct knowledge of the discussions. Microsoft ultimately wants capacity for more than 1 million chips, though component supplies and the TSMC negotiations could limit that plan.

Andrew Wall, general manager for Azure Maia, declined to discuss production volumes in a statement to The Information. Wall said the reported figures did not reflect the scale of Microsoft's program and that Microsoft expects Maia deployments to support AI demand measured in gigawatts. He did not provide a chip count or timetable.

The push follows a slow start for Maia 200. Microsoft delayed the accelerator after early tests missed internal targets, The Information reported, and Microsoft remains the only known user. Microsoft introduced Maia 200 in January as an inference chip built on TSMC's 3-nanometer process, with 216 gigabytes of HBM3e memory and support for lower-precision FP4 and FP8 calculations.

By April, Microsoft said Maia 200 was operating in its Iowa and Arizona data centers. Microsoft also claimed the accelerator delivered more than 30% better tokens per dollar than the newest silicon then running in Microsoft's fleet. That is an internal comparison, and Microsoft has not published enough customer testing to establish whether Maia can deliver similar savings across outside workloads.

Anthropic would give Maia an outside customer

Microsoft has held talks with Anthropic about using Maia, according to The Information. Landing Anthropic would give Microsoft the external validation that Maia 200 has yet to secure, while offering Anthropic another source of computing capacity for Claude.

Anthropic already operates at a scale that shows how far Microsoft has to go. AWS says nearly 1 million Trainium2 chips are training and serving Claude through Project Rainier. Alphabet separately told investors that Anthropic planned to access up to 1 million Google TPUs.

Those deployments give Amazon and Google something Microsoft still lacks: a major independent model developer building directly on their custom accelerators. Software tooling, model optimization and reliable access to large clusters can tie customers to a chip platform over successive hardware generations. Maia 300 therefore has to arrive with usable capacity and a mature software stack, rather than benchmark gains alone.

The reported order would remain smaller than rival programs. The Information cited Morgan Stanley estimates that Google planned to make more than 3 million TPUs in 2026 and 5 million in 2027. Microsoft's discussions for more than 300,000 Maia 300 chips would still represent an order-of-magnitude increase from Maia 200 production.

Microsoft can use Maia even without Anthropic

Microsoft has a large internal outlet for the chips. Maia runs OpenAI models and Microsoft's own MAI models, which support Copilot products, though The Information reported that Nvidia hardware still handles most of those workloads.

Microsoft's fallback is to direct more internal inference to Maia while reserving costly Nvidia systems for Azure customers willing to pay for them. Microsoft told investors in July that Maia 200 was 30% to 40% cheaper to operate than cutting-edge Nvidia chips when running OpenAI and Microsoft models, according to The Information. Internal testing has shown further gains from Maia 300, the report said, without providing a figure.

That division of capacity would let Microsoft capture savings on its own products and preserve scarce Nvidia hardware as a premium cloud resource. It would also give Microsoft's chip engineers a predictable workload while Azure tries to persuade outside developers to port models to Maia.

Microsoft has been pursuing this control since it introduced the first Maia accelerator in November 2023. A 2022 email later released in court showed Nadella worrying about Microsoft's thin infrastructure layer above Nvidia and its lack of control over both the chips and OpenAI's intellectual property, The Information reported.

Maia 300 is the first planned production ramp large enough to address that dependency. The September unveiling will establish Microsoft's technical claims. The consequential test comes in 2027, when Microsoft must secure the components, deploy hundreds of thousands of chips and convince at least one large customer to use them.

Reader comments

Conversation for this story loads after sign-in.