Ant Group Open-Sources Ming-Image-0.1-Design for Text-Heavy Graphics

Ant Group’s inclusionAI lab has open-sourced Ming-Image-0.1-Design, a pair of 6-billion-parameter image models built for graphic design rather than photography: user interfaces, infographics, posters, and other layouts where the rendered text has to be legible. The weights went up on Hugging Face on September 17, 2026, under the MIT licence, and the lab’s Ant Ling account announced the family on September 22. The release comes with a second model that turns a flat design back into editable transparent layers, plus two agent skills that build on both.

Intermediate

Three designs generated by Ming-Image-0.1-Design: a smartphone product banner, a retro clothing-store web page, and a virtual-reality game dashboard, each with dense rendered text
Sample outputs from Ming-Image-0.1-Design (cropped from the model card gallery). Image credit: inclusionAI, Ant Group

Two Models, One Design Workflow

The family has two checkpoints, each about 6.15 billion parameters:

  • Ming-Image-0.1-Design is a text-to-image model that inclusionAI describes as generating “complete visual compositions” for “UI, infographics, posters, and other text-rich visual designs.” It works from a prompt alone and takes no reference image. It can also output RGBA images with transparent backgrounds if the prompt starts with one of the recommended transparency phrases.
  • Ming-Image-0.1-Design-Layer goes the other way. Given a flattened design image and a layer plan, it returns the requested number of RGBA layers. Each layer is a separate PNG, and together they recompose the original.

Both run with 12 sampling steps in BF16. Text-to-image uses a CFG scale of 1.0 and produces square output at 1024 or 2048 pixels, with 2048 recommended. Layer decomposition uses a CFG scale of 2.0, works at a 512 or 1024 bucket and keeps the input’s aspect ratio. The lab says the “default and minimum validated deployment is one GPU with at least 80 GiB of memory,” so the release targets data-centre cards such as the A100 or H100 80 GB rather than consumer GPUs. For serving, inclusionAI recommends the vLLM-Omni inference framework.

Artificial Analysis UI/UX Design open-weights leaderboard showing Ming-Image-0.1-Design first at 1082 Elo, ahead of Ideogram 4.0 (Quality) at 1052 and FLUX.2 [dev] at 1000
The UI/UX Design leaderboard chart as published by inclusionAI. Image credit: inclusionAI, Ant Group (data: Artificial Analysis)

Prompts Written as Layouts

The model expects detailed prompts. inclusionAI publishes a rewriting step that runs before inference: an instruction-following vision-language model, either Ant’s Ling-3.0-flash-VL or Qwen3.8-27B, expands a short request into a structured JSON description. The repository describes it as “Figma-style layers ordered back to front, with exact coordinates, hierarchy, color specs, and every rendered string quoted verbatim and owned exactly once.” In practice, the prompt reads more like a layout spec than a caption, and each piece of text is spelled out exactly as it should appear.

The layer model uses a similar enhancer, with rules for how to split a design. Text goes on the front layer(s), and any card or banner behind the text becomes its own layer. The main subject gets a separate layer, and the background comes last, absorbing the surfaces and shadows underneath.

A greeting-card design decomposed into six transparent layers: title text, two sets of art supplies, a red ribbon, a black card panel, and a red background, followed by the recomposed result
Ming-Image-0.1-Design-Layer splitting a card design into six RGBA layers, then recomposing it. Image credit: inclusionAI, Ant Group

Benchmarks

The main claim is a ranking. In the chart inclusionAI published, Ming-Image-0.1-Design leads the open-weights view of Artificial Analysis’s UI/UX Design text-to-image leaderboard with an Elo of 1,082. Next come Ideogram 4.0 (Quality) at 1,052, Ideogram 4.0 at 1,015, and HunyuanImage 3.0 Instruct at 1,005. FLUX.2 [dev] scores 1,000, and Alibaba’s Z-Image Turbo, another 6B model, scores 946. Artificial Analysis computes these scores from blind preference votes. Two caveats apply: this is the open-weights view, not the overall board with proprietary models, and the figures come from inclusionAI’s screenshot rather than an independent evaluation.

For the layer model, inclusionAI reports results on the Crello test set of graphic designs. The table labels the method “CLEAR-1024.” With no layer merging allowed, it scores an RGB L1 error of 0.0574 (lower is better) and an alpha soft IoU of 0.8923 (higher is better). The open-source Qwen-Image-Layered at 1024 pixels scores 0.1409 and 0.7177. The lab’s table notes that a non-public version of Qwen-Image-Layered fine-tuned on Crello comes closer, at 0.0594 and 0.8705.

Table of layer-decomposition results on the Crello test set comparing Hi-SAM baselines, LayerD, Qwen-Image-Layered variants, and CLEAR-1024 on RGB L1 and alpha soft IoU
Layer-decomposition results on the Crello test set, as reported by inclusionAI. Image credit: inclusionAI, Ant Group

What This Means

Most open image models are tuned for photographs and illustrations, and text rendering is where they tend to break down. Ming-Image-0.1-Design specialises in the opposite case: dense, structured, text-heavy layouts. Pairing it with a decomposition model aims at a long-standing complaint about AI-generated design, which is that the output is one flat picture that is hard to edit afterwards.

The agent skills are the clearest sign of where Ant is heading. The Image-to-Editable-PPT skill, added to Ant’s ling-cookbook repository, takes a slide image, a screenshot, or an AI-rendered page and uses layer decomposition to rebuild it as a native PowerPoint slide. Ordinary text becomes text boxes, simple containers become shapes, and complex artwork becomes separate cropped images. The Ling UI Design Skill applies the same idea to interface design. Both skills are aimed at editable files rather than finished images.

Practical limits remain. The 80 GiB validated configuration puts local use out of reach for most people until quantised builds appear. The prompt-rewriting step adds a second model to the pipeline. And the “0.1” version number suggests that inclusionAI considers this an early release.

Related Coverage

This post was drafted with AI assistance and reviewed by RITS staff.

Sources