Single-stream DiT with Prompt Enhancer
According to the model card, the model combines a single-stream Diffusion Transformer with a lightweight Prompt Enhancer for richer structured descriptions.
Open Source Model Profile · baidu
ERNIE-Image is an 8.03B-parameter text-to-image model from Baidu. According to the model card, it pairs a single-stream Diffusion Transformer with a Prompt Enhancer for instruction following and text rendering.
ERNIE-Image is published by Baidu as a text-to-image model. Safetensors metadata reports 8,033,490,048 parameters, and hub metadata records Diffusers support with an Apache-2.0 license. According to the model card, it is built on a single-stream Diffusion Transformer with a lightweight Prompt Enhancer that expands brief inputs into richer descriptions.
According to the model card, the model combines a single-stream Diffusion Transformer with a lightweight Prompt Enhancer for richer structured descriptions.
According to the model card, it is presented as strong on dense, long-form, and layout-sensitive text for posters, infographics, and UI-like images.
According to the model card, the SFT release targets general capability in typically 50 steps, while the Turbo release uses DMD and RL for 8-step generation.
According to the model card, recommended settings include resolutions such as 1024x1024 and 1264x848, guidance scale 4.0, and ErnieImagePipeline usage.
Source: baidu/ERNIE-Image
Captured: Unknown. Processed: 2026-09-07T19:35:54.450614+00:00.
ERNIE-Image 🤗 ERNIE-Image | 🤗 ERNIE-Image-Turbo | 🤖 ERNIE-Image | 🤖 ERNIE-Image-Turbo 🖥️ Huggingface Demo1 | 🖥️ Huggingface Demo2(ZeroGPU) | 🖥️ AI Studio Demo Github | 📖 Blog | 🖼️ Art Gallery 💬 WeChat(微信) | 🫨 Discord | 🏷️ X ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT) and paired with a lightweight Prompt Enhancer that expands brief user inputs into richer structured descriptions. With only 8B DiT parameters, it reaches state-of-the-art performance among open-weight text-to-image models. The model is designed no…
F001F002F003F004F005F006F007F008F009F010F011F013F014F015F016F017F018F019