Skip to content

EthenEthenEthen

Open Source Model Profile · Tongyi-MAI

Z-Image-Turbo

Z-Image-Turbo is a 6.15B-parameter distilled text-to-image model from Tongyi-MAI. According to the model card, it is the few-step distilled variant of Z-Image operating at 8 function evaluations.

Publisher
Tongyi-MAI
Task
text-to-image
Model type
Unknown
License
apache-2.0
Library
diffusers
Publication status
Accepted · not indexed

Model overview

Z-Image-Turbo is published by Tongyi-MAI as a text-to-image generation model. Captured Safetensors metadata reports 6,154,908,736 parameters, or about 6.15B. According to the model card, it is a distilled version of the Z-Image foundation model with a single-stream diffusion transformer design, supporting photorealistic generation, English and Chinese text rendering, and instruction adherence.

Recorded capabilities

Distilled few-step generation

According to the model card, the Turbo variant is distilled to run at 8 function evaluations.

Single-stream DiT design

According to the model card, the architecture concatenates text, visual semantic tokens, and image VAE tokens into one unified input stream.

Publisher-reported latency and VRAM

According to the model card, the model offers sub-second inference latency on H800 GPUs and fits within 16GB VRAM consumer devices.

Diffusers integration

According to the model card, two diffusers pull requests adding Z-Image support were merged, with ZImagePipeline loadable in bfloat16.

Use cases in the source record

  • Few-step text-to-image generation using the model card's ZImagePipeline workflow at 1024x1024 with 9 inference steps and zero guidance.
  • Bilingual image-text rendering experiments in English and Chinese as described in the model card.
  • Downstream fine-tuning and creative generation starting from the Z-Image foundation lineage described in the card.

Limitations and unknowns

  • No independently measured evaluation results were extracted; quality comparisons are publisher-reported card statements.
  • Latency and VRAM figures come from the publisher model card and were not independently verified.
  • No context-window or tokenizer detail was extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: Tongyi-MAI/Z-Image-Turbo

Captured: Unknown. Processed: 2026-09-07T19:34:38.146002+00:00.

⚡️- Image An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Welcome to the official repository for the Z-Image(造相)project! ✨ Z-Image Z-Image is a powerful and highly efficient image generation model family with 6B parameters. Currently there are four variants: 🚀 Z-Image-Turbo – A distilled version of Z-Image that matches or exceeds leading competitors with only 8 NFEs (Number of Function Evaluations). It offers ⚡️sub-second inference latency⚡️ on enterprise-grade H800 GPUs and fits comfortably within 16G VRAM consumer devices . It excels in photorealistic image generation, bilingual text render…

F001F002F003F004F005F006F008F009F010F011F015F016F018F019