Microsoft ships MAI-Image-2.6 to Foundry at about 4 cents an image
Mustafa Suleyman's lab paired Arena's No. 2 text-to-image model with a half-price Flash version for high-volume workloads.
By Ryan Merket · Published
Primary source: Arena on X
Why it matters
Microsoft now has a top-tier image model it can sell and deploy itself, giving Suleyman's lab more control over inference costs and less dependence on outside model providers.

Mustafa Suleyman (@mustafasuleyman) put Microsoft AI's highest-ranked image model into public preview on Friday, turning a month-old leaderboard result into a product developers can deploy through Microsoft Foundry.
MAI-Image-2.6 is accompanied by MAI-Image-2.6-Flash, a cheaper version designed for latency-sensitive and high-volume applications. Both models support generation and editing with multiple reference images, web grounding, dynamic aspect ratios and resolutions up to 1.5K.
Suleyman, the DeepMind and Inflection co-founder who formed Microsoft AI after joining Microsoft in 2024, is building the company's own model portfolio alongside its continuing distribution of models from outside labs. During a March 2026 restructuring, he wrote that the industry would be defined by "frontier models, and the products through which they are experienced." MAI-Image-2.6 advances both sides of that mandate: Microsoft owns the model and controls its route into Foundry and, eventually, Microsoft's consumer products.
Arena said in a post on X that MAI-Image-2.6 scored 1,331 points and ranked No. 2 in its Text-to-Image Arena, behind OpenAI's GPT Image 2. Arena also placed Microsoft's model second in its product, branding and commercial design category.
The live Text-to-Image Arena leaderboard subsequently showed MAI-Image-2.6 at 1,332 points, with a seven-point confidence interval. GPT Image 2 led at 1,382, while xAI's preliminary Grok Imagine Image 2.0 score stood at 1,315. The changing score reflects the continuous flow of user votes rather than a fixed benchmark result.
Arena builds the rankings from blind, head-to-head comparisons. Users enter a prompt, receive images from two unidentified models and vote for the better output. That method captures human preference across real prompts, although it does not measure production reliability, latency or consistency under a developer's particular workload.
The model's price is central to Microsoft's pitch. Arena estimated MAI-Image-2.6 at $38.90 per 1,000 output images, or roughly $0.039 each, and placed it on the service's price-performance Pareto frontier. In practical terms, Arena found no other evaluated model that was both cheaper and more highly rated under its assumptions. Arena said Microsoft's model outscored Grok Imagine Image 2.0 while costing less.
Microsoft bills developers by tokens rather than by a flat image fee. MAI-Image-2.6 starts at $5 per 1 million text input tokens, $8 per 1 million image input tokens and $38 per 1 million image output tokens. MAI-Image-2.6-Flash cuts those rates to $1.75, $2.50 and $19, respectively. Actual per-image spending will depend on resolution, inputs and the number of output tokens consumed, so Arena's four-cent figure should be treated as an estimate rather than a universal invoice price.
Microsoft says the Flash model produces images 2.8 times faster than GPT Image 2 Medium while using resources 72% more efficiently. Those comparative figures come from Microsoft, and the company has not published enough detail in the launch post to translate them across every workload or serving configuration.
The September 4th release is the availability event, rather than the model's first appearance. Microsoft announced MAI-Image-2.6 on August 10th after it entered Arena testing. By August 18th, the model was available in Microsoft's playground and in a private Foundry preview. Friday's launch opened the Foundry preview to the public and introduced the Flash variant.
Microsoft has been explicit about the business logic behind its first-party models. On its fiscal 2026 fourth-quarter earnings call, the company said it was designing the MAI family around cost-efficient inference and a system in which models can be substituted according to quality, latency, price and compliance requirements. That gives Microsoft an internal option for workloads where paying another model provider would compress margins, while Foundry can continue selling customers access to competing models.
MAI-Image-2.6 still trails OpenAI's model by about 50 Arena points, and its small lead over Grok sits within overlapping statistical ranges. The useful result for Microsoft is commercial: Suleyman's lab has reached the top tier of crowd-ranked image generators with a model priced for repeated production use. Foundry customers can now test whether that leaderboard position survives contact with their own prompts, brand rules and cloud bills.