Microsoft is making another push into the increasingly crowded AI image-generation market with MAI-Image-2.6, its latest image model designed to deliver higher-quality visuals while keeping generation costs under control.
The new model is now available to developers through Microsoft Foundry in public preview, alongside a faster MAI-Image-2.6-Flash variant aimed at production workloads where speed and efficiency matter as much as image quality. Microsoft says both models are built around the practical needs of creators and businesses, rather than simply producing impressive one-off demos.
MAI-Image-2.6 focuses on quality and control
Microsoft describes MAI-Image-2.6 as its strongest image model yet. The model supports both text-to-image generation and image editing, with capabilities designed to give users more control over composition, style, references and visual details.
One of the biggest additions is multi-reference editing. Users can combine multiple reference images in a single request, allowing a product, person, character, brand element or visual style to be carried across a new creation. Microsoft says the model can accept up to five reference images, which could make it particularly useful for commercial creative workflows where visual consistency matters.
The model also supports web grounding. Rather than relying exclusively on information contained within a prompt, MAI-Image-2.6 can pull relevant information from the web to help create visuals based on current context and documents. That could be useful for everything from marketing materials and product imagery to visuals that need to reflect up-to-date information.
Microsoft has also added dynamic aspect ratios and support for resolutions of up to 1.5K. The idea is to make the model better suited to different creative formats instead of forcing users into a single fixed image shape.
The improvements aren’t limited to composition. Microsoft says MAI-Image-2.6 delivers stronger text rendering, better portraits and 3D imagery, as well as more polished commercial and photorealistic results. These improvements are particularly relevant for product photography, branding, advertising and cinematic imagery, where inaccurate text or inconsistent objects can quickly make generated images unusable.
A strong showing on image-generation benchmarks
MAI-Image-2.6 is also arriving with some notable benchmark results.
According to Microsoft, the model ranks No. 2 for both text-to-image generation and image editing on Arena as of September 4. On Artificial Analysis, Microsoft says it ranks No. 2 for text-to-image and No. 1 for image editing.
That represents a substantial improvement over MAI-Image-2.5. When Microsoft first announced MAI-Image-2.6 in August, it said the model had gained 79 Elo points over its predecessor across Arena’s text-to-image evaluation, with text rendering alone improving by 91 points.
Those results put Microsoft’s model firmly into competition with the image-generation systems coming from other major AI companies. Microsoft previously highlighted MAI-Image-2.6’s performance against models from Google, Meta and SpaceXAI, signaling how seriously the company is taking the generative-image market.
MAI-Image-2.6-Flash is built for speed
Alongside the flagship model, Microsoft is introducing MAI-Image-2.6-Flash.
The Flash model is designed for high-throughput and latency-sensitive applications where developers may need to generate large numbers of images quickly. Microsoft says it can generate images 2.8 times faster than GPT-Image-2-Medium while delivering 72% greater efficiency.
The company positions the two models differently: MAI-Image-2.6 is aimed at maximum precision and quality, while the Flash version is intended for production environments where faster generation and lower costs can have a significant impact at scale.
Microsoft’s Foundry documentation confirms that both models support text-to-image generation and image-to-image editing and are currently available as preview models.
Pricing also reflects that distinction. Microsoft lists MAI-Image-2.6 at $5 per 1 million text-input tokens, $8 per 1 million image-input tokens and $38 per 1 million image-output tokens. MAI-Image-2.6-Flash starts at $1.75 per 1 million text-input tokens, $2.50 per 1 million image-input tokens and $19 per 1 million image-output tokens.
That price gap is important because Microsoft isn’t simply competing on benchmark scores. The company is explicitly pitching MAI-Image-2.6 around the relationship between quality and cost, claiming the model offers the industry’s best price-per-Elo performance.
For businesses generating thousands or even millions of images, that equation could ultimately matter more than winning an individual benchmark.
Microsoft is targeting real creative workloads
The broader strategy behind MAI-Image-2.6 is becoming clearer as Microsoft’s MAI model family expands.
The company isn’t presenting image generation as just a novelty for consumers. Its focus is increasingly on professional workflows where generated images need to be consistent, editable and usable across multiple formats.
That explains the emphasis on reference images, accurate text, web grounding, flexible aspect ratios and commercial imagery. A marketing team, retailer or product designer needs an image model that can follow a creative brief and preserve important visual details, not just produce something attractive from a short prompt.
MAI-Image-2.6 is now available through Microsoft Foundry in public preview, while users can also try the models through Microsoft’s MAI Playground.
With MAI-Image-2.6 and the cheaper, faster Flash variant, Microsoft is betting that the next stage of AI image generation won’t be defined solely by which model produces the prettiest picture. It will be about how reliably that model can produce the right picture, at the right speed and at a cost businesses can actually afford.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
