Skip to content

EthenEthenEthen

Open Source Model Profile · stabilityai

stable-diffusion-3-medium

stable-diffusion-3-medium is a Stability AI Multimodal Diffusion Transformer for text-to-image generation. Its model card documents triple text encoders and three weight-bundle variants.

Publisher
stabilityai
Task
text-to-image
Model type
Unknown
License
stabilityai-ai-community
Library
diffusion-single-file
Publication status
Accepted · not indexed

Model overview

stable-diffusion-3-medium is published by Stability AI as a text-to-image generative model. According to the model card, it is a Multimodal Diffusion Transformer for generating images from text prompts. Card data records other, with the card specifying the Stability Community License.

Recorded capabilities

MMDiT text-to-image model

According to the model card, this is a Multimodal Diffusion Transformer for text-prompt image generation with improved image quality, typography, and prompt understanding.

Triple text-encoder design

The model card says it uses three fixed pretrained text encoders: OpenCLIP-ViT/G, CLIP-ViT/L, and T5-xxl.

Three documented weight bundles

According to the model card, the three safetensors bundles share MMDiT and VAE weights and differ only in text-encoder inclusion and precision.

ComfyUI and Diffusers paths

The model card recommends ComfyUI for self-hosted use and publishes separate diffusers-compatible weights with a StableDiffusion3Pipeline example.

Use cases in the source record

  • Text-prompt image generation, including complex prompts where the publisher notes typography and prompt-understanding behavior.
  • Self-hosted generation through ComfyUI or the diffusers StableDiffusion3Pipeline path, following the publisher's documented setups.

Limitations and unknowns

  • No parameter count was extracted from this record.
  • No evaluation scores were extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • Safety, factuality, and misuse limits are publisher-stated; the card asks developers to test and apply guardrails for their use case.

Source and provenance

Source: stabilityai/stable-diffusion-3-medium

Captured: Unknown. Processed: 2026-09-07T19:34:58.564051+00:00.

Stable Diffusion 3 Medium Model Stable Diffusion 3 Medium is a Multimodal Diffusion Transformer (MMDiT) text-to-image model that features greatly improved performance in image quality, typography, complex prompt understanding, and resource-efficiency. For more technical details, please refer to the Research paper . Please note: this model is released under the Stability Community License. For Enterprise License visit Stability.ai or contact us for commercial licensing details. Model Description Developed by: Stability AI Model type: MMDiT text-to-image generative model Model Description: This is a model that can be used to generate…

F001F002F003F004F007F008F010F011F012F013F014F015F016F017F018F019F021F022F023