Skip to content

EthenEthenEthen

Open Source Model Profile · baidu

ERNIE-Image

ERNIE-Image is an 8.03B-parameter text-to-image model from Baidu. According to the model card, it pairs a single-stream Diffusion Transformer with a Prompt Enhancer for instruction following and text rendering.

Publisher
baidu
Task
text-to-image
Model type
Unknown
License
apache-2.0
Library
diffusers
Publication status
Approved for indexing

Model overview

ERNIE-Image is published by Baidu as a text-to-image model. Safetensors metadata reports 8,033,490,048 parameters, and hub metadata records Diffusers support with an Apache-2.0 license. According to the model card, it is built on a single-stream Diffusion Transformer with a lightweight Prompt Enhancer that expands brief inputs into richer descriptions.

Recorded capabilities

Single-stream DiT with Prompt Enhancer

According to the model card, the model combines a single-stream Diffusion Transformer with a lightweight Prompt Enhancer for richer structured descriptions.

Documented text-rendering focus

According to the model card, it is presented as strong on dense, long-form, and layout-sensitive text for posters, infographics, and UI-like images.

SFT and Turbo releases

According to the model card, the SFT release targets general capability in typically 50 steps, while the Turbo release uses DMD and RL for 8-step generation.

Diffusers deployment detail

According to the model card, recommended settings include resolutions such as 1024x1024 and 1264x848, guidance scale 4.0, and ErnieImagePipeline usage.

Use cases in the source record

  • Text-to-image generation for posters, comics, multi-panel layouts, infographics, and other layout-sensitive visual content, as described in the model card.
  • Diffusers and SGLang generation workflows using the documented ErnieImagePipeline, resolutions, 50-step SFT defaults, and prompt-enhancer option.

Limitations and unknowns

  • State-of-the-art and benchmark-table figures are publisher claims and were not independently verified.
  • No context-window value was extracted from this record.
  • Provider state is historical snapshot data showing an error model status, not independently refreshed current availability.

Source and provenance

Source: baidu/ERNIE-Image

Captured: Unknown. Processed: 2026-09-07T19:35:54.450614+00:00.

ERNIE-Image 🤗 ERNIE-Image | 🤗 ERNIE-Image-Turbo | 🤖 ERNIE-Image | 🤖 ERNIE-Image-Turbo 🖥️ Huggingface Demo1 | 🖥️ Huggingface Demo2(ZeroGPU) | 🖥️ AI Studio Demo Github | 📖 Blog | 🖼️ Art Gallery 💬 WeChat(微信) | 🫨 Discord | 🏷️ X ERNIE-Image is an open text-to-image generation model developed by the ERNIE-Image team at Baidu. It is built on a single-stream Diffusion Transformer (DiT) and paired with a lightweight Prompt Enhancer that expands brief user inputs into richer structured descriptions. With only 8B DiT parameters, it reaches state-of-the-art performance among open-weight text-to-image models. The model is designed no…

F001F002F003F004F005F006F007F008F009F010F011F013F014F015F016F017F018F019